Specifications
Model
Taalas HC1
Product Type
Hardwired (mask-programmed) LLM inference accelerator
Architecture
Model-specific silicon — Llama-3.1 8B implemented directly in hardware
Throughput
Approximately 17,000 tokens/s on typical prompts; up to ~20,000 tokens/s on short queries
Demonstrated Sample
14-chapter outline generated in 0.064 s at 15,651 tokens/s
Process Node
TSMC 6nm
Die Area
815 mm²
Transistors
53 billion
Memory Architecture
Storage and compute unified on a single chip at DRAM-level density
Bottleneck Removed
Avoids the memory-bandwidth bottleneck that limits conventional accelerators on LLM workloads
Model Flexibility
Configurable context window size; fine-tuning via low-rank adapters (LoRAs)
Model Coverage
Llama-3.1 8B at launch; a mid-sized reasoning LLM on the same HC1 silicon follows in Q2
Power Class
Designed for 2.5 kW servers
Vendor Efficiency Claims
Roughly 10x faster than a Cerebras chip on the same model, around 20x lower build cost and 10x lower power (Taalas claims)
Ecosystem Position
Targets multiuser server inference and latency-critical interactive workloads such as voice interaction
Roadmap
Second-generation HC2 silicon for higher density and faster execution, with deployments starting by end of year
Runtime
Vendor-supplied stack with a public online chatbot demonstration for evaluation
Overview
The Taalas HC1 takes an unusual position in the AI accelerator market: instead of a general-purpose matrix engine, it is hardwired — physically implemented in silicon — for Llama-3.1 8B. The chip delivers close to 17,000 tokens/s running that model, with the shortest queries reaching roughly 20,000 tokens/s.
The design rationale is memory. Conventional accelerators place memory on one side and compute on the other, and because the two run at different speeds, memory bandwidth becomes the bottleneck for large language models. Taalas unifies storage and compute on a single chip at DRAM-level density, which raises throughput and drops power. The device is manufactured on TSMC's 6nm process, measures 815 mm² and carries 53 billion transistors, and is designed for 2.5 kW servers.
The trade-off is fixed model coverage. HC1 currently runs Llama-3.1 8B, and there is no way to load an arbitrary model; Taalas states that flexibility is retained through configurable context window size and fine-tuning via low-rank adapters. The vendor claims roughly 10x the throughput of a Cerebras chip on the same model at 20x lower build cost and 10x lower power. Because the current implementation targets 2.5 kW servers, robotics voice interaction — one of the motivating use cases — remains a later-generation target; a mid-sized reasoning LLM on the same silicon is scheduled for Q2, with the higher-density HC2 platform following by year end.
Key Benefits
Close to 17,000 tokens/s from a single accelerator and up to 20,000 tokens/s on short queries. Unified storage and compute on one die at DRAM-level density removes the memory-bandwidth bottleneck that caps conventional LLM accelerators. 53 billion transistors on 815 mm² of TSMC 6nm in a 2.5 kW server class package. Configurable context window and LoRA fine-tuning support retain a measure of deployment flexibility. A defined roadmap to HC2 signals a path to higher density and faster execution.
Applications
High-throughput multiuser LLM serving, latency-sensitive conversational AI, text generation appliances, accelerator research and benchmarking against general-purpose datacenter accelerators, model-specific inference where throughput per watt dominates.
Request a Quote — TAALAS HC1 — HARDWIRED LLAMA-3.1 8B INFERENCE ACCELERATOR AT ~17,000 TOKENS/S
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →