Specifications

Architecture

Fixed-function transformer ASIC — transformer operations implemented directly in hardware

Memory per Chip

144 GB HBM3E

Memory Bandwidth

Roughly 1.8x the HBM bandwidth of NVIDIA H100 SXM5

Programmability

None — hardwired operations only, with no general-purpose programmability

Software Stack

Not vLLM-compatible; requires Etched's own serving stack

Process Node

4 nm

Single-Chip Throughput

Approximately 62,000 tokens/sec quoted for Llama 70B at batch 1, derived from an 8-chip server figure of 500,000 tok/s

Comparison Point

Roughly 700 tok/s for Llama 70B at batch 1 on H100 SXM5 and about 1,200 tok/s on B200 SXM6 under vLLM

Server Configuration

Rack-scale systems built around eight Sohu chips

Design Rationale

Transformer inference is the dominant computational bottleneck, so specialisation trades architectural flexibility for throughput and cost

Memory Contracts

Memory supply sourced through Rambus; wafers booked at TSMC

Named Customer

Jane Street, the quantitative trading firm, reported as an early customer

Commercial Status

Rack-scale systems; the first rack-scale system was reported shipped during 2026

Availability Caveat

Pre-volume — public sources during 2026 disagreed on whether production units had shipped in volume to paying customers, so treat broad availability as emerging

Company

San Jose-based startup founded by Harvard dropouts, with a reported 2026 Series C co-led by Sequoia and Andreessen Horowitz

Overview

Etched Sohu takes the opposite bet to general-purpose GPUs: instead of a programmable processor that can run anything, it hardwires the transformer architecture into silicon. That removes the instruction-fetch and scheduling overheads a GPU carries for every operation, and lets the memory subsystem be sized for one workload — transformer inference — rather than for generality.

Each Sohu chip pairs roughly 1.8x the HBM bandwidth of an NVIDIA H100 SXM5 with 144 GB of HBM3E on a 4 nm process. The trade-off is explicit: there is no vLLM compatibility and no general programmability, so deployments run Etched's own serving stack. Etched's published figures describe an eight-chip system serving around 500,000 tokens per second for Llama 70B, which works out to roughly 62,000 tok/s per chip at batch 1.

The commercial picture is genuinely early. A rack-scale system was reported shipped during 2026 with Jane Street named as an early customer, but public reporting through the year disagreed on whether volume production silicon had reached paying customers. Buyers evaluating transformer-specific silicon should weigh the throughput-per-dollar case against model-architecture risk: a chip that implements only transformers is exposed if architectures shift toward state-space or hybrid designs. QS Compute supplies LLM inference accelerators and single-purpose AI silicon — contact us to discuss fit against your serving workload.

Key Benefits

Throughput per watt: hardwired transformer datapath removes GPU scheduling overhead. Memory-dense: 144 GB HBM3E per chip with roughly 1.8x H100 bandwidth. Cost per token: specialisation targets substantially lower inference cost for high-volume LLM serving. Rack-scale delivery: turnkey eight-chip systems rather than bare accelerators.

Applications

High-volume LLM inference serving, chat and assistant endpoints, batch text generation, retrieval-augmented serving pipelines, and cost-sensitive deployments where the served model family is stable and known.

Request a Quote — ETCHED SOHU — TRANSFORMER-ONLY INFERENCE ASIC

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

NVIDIA Groq 3 LPU Cerebras WSE-3 Turbo SambaNova SN40L