3nm TDN AIP processor built on logarithmic math · 72-chip TDN72 air-cooled pod at 76.7 PFLOPS FP16 / 30kW · four-pod rack at 306 PFLOPS FP16 / 120kW · NapierLink under 1,000ns at 1 TB/s any-to-any
Tensordyne's 3nm AI inference processor integrating the Universal Inference Engine with balanced HBM and SRAM.
A 72-chip air-cooled inference pod built around a 72-node scale-up interconnect.
Four independent TDN72 pods unified into a single inference rack.
Tensordyne's any-to-any fabric that unifies 72 chips into what behaves as one large accelerator.
Tensordyne (formerly Recogni)
Tensordyne Napier Inference System
TDN AIP (Tensordyne Napier AI Processor)
TSMC 3nm
Logarithmic math — multiplication converted to addition
Balanced HBM and SRAM on the Universal Inference Engine
TDN72 — 72 chips, 76.7 PFLOPS FP16, 30 kW
4 × TDN72 — 306 PFLOPS FP16, 120 kW
NapierLink — under 1,000ns, 1 TB/s any-to-any
Fully air cooled
Standard token serving and disaggregated inference (software toggle)
Logarithmic quantization — 16-bit models at a fraction of compute, zero accuracy loss
Python-inspired eDSL with Torch-like interface, Hugging Face model zoo
Broadcom IP platform, HPE Juniper Networks, TSMC manufacturing
3nm silicon taped out; shipping via early partners
Available
Quote
Tensordyne Napier is a rack-scale AI inference system from Tensordyne — the company formerly known as Recogni, which pivoted from autonomous-vehicle silicon to AI inference after being founded in 2017. Napier is built on a core architectural bet: replacing the multiplication that dominates AI matrix math with logarithmic addition, embedded in custom silicon and unified by a very low-latency interconnect.
The system's processor, the TDN AIP, is manufactured by TSMC on a 3nm node in partnership with Broadcom (IP platform, packaging and silicon design) and HPE Juniper Networks. It combines a Universal Inference Engine with balanced HBM and SRAM, and bakes interconnect logic directly into the die. Tensordyne says the shift to addition yields reduced power and area, which it spends on wider data paths and more on-chip memory.
Napier scales out through the TDN72 — a 72-chip air-cooled pod delivering 76.7 PFLOPS of FP16 compute within a 30 kW envelope — and the full Napier rack, which unifies four pods for roughly 306 PFLOPS FP16 at 120 kW. A NapierLink fabric delivers under 1,000 nanoseconds of latency and 1 TB/s any-to-any bandwidth, which Tensordyne says enables near-linear scaling across 72 chips and support for 10-trillion-parameter mixture-of-experts models.
Tensordyne reports 363,000 tokens per second per rack on DeepSeek-R1 against 27,400 for an NVIDIA NVL72 GB300, 11 million tokens per kWh, and claims of 9× better space efficiency, 2× the speed and 10× cost savings versus the NVIDIA-plus-Groq standard reference. The platform is entirely air cooled, ships with a Python-inspired eDSL and Torch-like interface, and supports both standard token serving and disaggregated inference via a software toggle. Tensordyne reports more than a dozen Letters of Intent worth over $200M, with Cirrascale an announced customer.
Need Tensordyne Napier?
Contact QS Compute for availability, configuration, and volume pricing.
Request Quote