Specifications
Matrix Compute
13.4 PFLOP/s of MXFP4 x MXFP4 matrix compute per chip
Memory per Chip
216 GiB of HBM4
Memory Bandwidth
15.4 TB/s of HBM4 bandwidth
Package Power
700 W per accelerator
Maximum System Size
2,048 chips per system
Aggregate System Compute
27 EFLOP/s
Aggregate System Memory
432 TiB of HBM across a full system
Scale-up Network
600 GB/s within the local 128-ASIC domain
Scale-out Network
200 GB/s across the global 2,048-ASIC domain
Programming Model
Spatial programming model that compiles functional kernels into high-performance implementations
Design Cycle
Approximately nine months from initial RTL to tapeout
Process Node
TSMC 3 nm
Build Partners
OpenAI accelerator architecture, Broadcom silicon implementation plus networking and connectivity, Celestica boards, racks and systems
Performance Positioning
Positioned against NVIDIA GB200 and GB300; OpenAI claims roughly 50% lower inference cost per token than current-generation GPUs
Deployment
Initial deployment targeted for the end of 2026, beginning through Microsoft Azure and expanding to other datacentre partners
Roadmap
Gen 1 of a multi-generation platform — Gen 2 targets better performance per watt, Gen 3 targets economical low-latency serving
Announced
Unveiled June 24, 2026; architecture detailed at Hot Chips 2026
Overview
Jalapeño is OpenAI's first Intelligence Processor: an inference accelerator architected around OpenAI's own production understanding of how large language models behave — attention patterns, kernel profiles and memory access habits at planetary scale. Broadcom handled the silicon implementation plus networking and connectivity, and Celestica contributed the board, rack and system engineering.
The architecture is built around HBM4 and a spatial programming model rather than a conventional general-purpose GPU pipeline. Each chip delivers 13.4 PFLOP/s of MXFP4 matrix compute against 216 GiB of HBM4 running at 15.4 TB/s, inside a 700 W package. Scaling is explicitly two-tier: a tightly coupled 128-ASIC domain at 600 GB/s, and a global 2,048-ASIC domain at 200 GB/s, which together form a 27 EFLOP/s, 432 TiB system.
Jalapeño is Gen 1 of a stated multi-generation platform. OpenAI's engineers took the design from initial RTL to tapeout in roughly nine months, and the company has described the next generations as targeting performance per watt first and then economical low-latency serving, with the goal of unlocking aggregate HBM bandwidth rather than simply adding raw bandwidth. Initial deployment is targeted for the end of 2026 starting through Microsoft Azure. QS Compute supplies LLM inference accelerators, GPU servers and rack-scale AI systems — request a quote.
Key Benefits
Inference-first silicon: MXFP4 matrix engines sized for LLM serving rather than training. Huge fast memory: 216 GiB HBM4 at 15.4 TB/s per chip. Rack-scale fabric: 600 GB/s local and 200 GB/s global domains across 2,048 chips. Rapid iteration: nine months from RTL to tapeout, with Gen 2 and Gen 3 already scoped.
Applications
Large-language-model inference at datacentre scale, long-context and retrieval-augmented serving, agentic AI workloads, chat and assistant platforms, and high-throughput token generation where cost per token dominates the deployment decision.
Request a Quote — OPENAI JALAPEÑO — BROADCOM-BUILT LLM INFERENCE ASIC
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →