Specifications

Model

Taalas HC1

Product Type

Hardwired (mask-programmed) LLM inference accelerator

Architecture

Model-specific silicon — Llama-3.1 8B implemented directly in hardware

Throughput

Approximately 17,000 tokens/s on typical prompts; up to ~20,000 tokens/s on short queries

Demonstrated Sample

14-chapter outline generated in 0.064 s at 15,651 tokens/s

Process Node

TSMC 6nm

Die Area

815 mm²

Transistors

53 billion

Memory Architecture

Storage and compute unified on a single chip at DRAM-level density

Bottleneck Removed

Avoids the memory-bandwidth bottleneck that limits conventional accelerators on LLM workloads

Model Flexibility

Configurable context window size; fine-tuning via low-rank adapters (LoRAs)

Model Coverage

Llama-3.1 8B at launch; a mid-sized reasoning LLM on the same HC1 silicon follows in Q2

Power Class

Designed for 2.5 kW servers

Vendor Efficiency Claims

Roughly 10x faster than a Cerebras chip on the same model, around 20x lower build cost and 10x lower power (Taalas claims)

Ecosystem Position

Targets multiuser server inference and latency-critical interactive workloads such as voice interaction

Roadmap

Second-generation HC2 silicon for higher density and faster execution, with deployments starting by end of year

Runtime

Vendor-supplied stack with a public online chatbot demonstration for evaluation

Overview

The Taalas HC1 takes an unusual position in the AI accelerator market: instead of a general-purpose matrix engine, it is hardwired — physically implemented in silicon — for Llama-3.1 8B. The chip delivers close to 17,000 tokens/s running that model, with the shortest queries reaching roughly 20,000 tokens/s.

The design rationale is memory. Conventional accelerators place memory on one side and compute on the other, and because the two run at different speeds, memory bandwidth becomes the bottleneck for large language models. Taalas unifies storage and compute on a single chip at DRAM-level density, which raises throughput and drops power. The device is manufactured on TSMC's 6nm process, measures 815 mm² and carries 53 billion transistors, and is designed for 2.5 kW servers.

The trade-off is fixed model coverage. HC1 currently runs Llama-3.1 8B, and there is no way to load an arbitrary model; Taalas states that flexibility is retained through configurable context window size and fine-tuning via low-rank adapters. The vendor claims roughly 10x the throughput of a Cerebras chip on the same model at 20x lower build cost and 10x lower power. Because the current implementation targets 2.5 kW servers, robotics voice interaction — one of the motivating use cases — remains a later-generation target; a mid-sized reasoning LLM on the same silicon is scheduled for Q2, with the higher-density HC2 platform following by year end.

Key Benefits

Close to 17,000 tokens/s from a single accelerator and up to 20,000 tokens/s on short queries. Unified storage and compute on one die at DRAM-level density removes the memory-bandwidth bottleneck that caps conventional LLM accelerators. 53 billion transistors on 815 mm² of TSMC 6nm in a 2.5 kW server class package. Configurable context window and LoRA fine-tuning support retain a measure of deployment flexibility. A defined roadmap to HC2 signals a path to higher density and faster execution.

Applications

High-throughput multiuser LLM serving, latency-sensitive conversational AI, text generation appliances, accelerator research and benchmarking against general-purpose datacenter accelerators, model-specific inference where throughput per watt dominates.

Request a Quote — TAALAS HC1 — HARDWIRED LLAMA-3.1 8B INFERENCE ACCELERATOR AT ~17,000 TOKENS/S

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

NVIDIA Groq 3 LPU Tenstorrent Blackhole P150 Cambricon MLU370-X8