Specifications

Matrix Compute

13.4 PFLOP/s of MXFP4 x MXFP4 matrix compute per chip

Memory per Chip

216 GiB of HBM4

Memory Bandwidth

15.4 TB/s of HBM4 bandwidth

Package Power

700 W per accelerator

Maximum System Size

2,048 chips per system

Aggregate System Compute

27 EFLOP/s

Aggregate System Memory

432 TiB of HBM across a full system

Scale-up Network

600 GB/s within the local 128-ASIC domain

Scale-out Network

200 GB/s across the global 2,048-ASIC domain

Programming Model

Spatial programming model that compiles functional kernels into high-performance implementations

Design Cycle

Approximately nine months from initial RTL to tapeout

Process Node

TSMC 3 nm

Build Partners

OpenAI accelerator architecture, Broadcom silicon implementation plus networking and connectivity, Celestica boards, racks and systems

Performance Positioning

Positioned against NVIDIA GB200 and GB300; OpenAI claims roughly 50% lower inference cost per token than current-generation GPUs

Deployment

Initial deployment targeted for the end of 2026, beginning through Microsoft Azure and expanding to other datacentre partners

Roadmap

Gen 1 of a multi-generation platform — Gen 2 targets better performance per watt, Gen 3 targets economical low-latency serving

Announced

Unveiled June 24, 2026; architecture detailed at Hot Chips 2026

Overview

Jalapeño is OpenAI's first Intelligence Processor: an inference accelerator architected around OpenAI's own production understanding of how large language models behave — attention patterns, kernel profiles and memory access habits at planetary scale. Broadcom handled the silicon implementation plus networking and connectivity, and Celestica contributed the board, rack and system engineering.

The architecture is built around HBM4 and a spatial programming model rather than a conventional general-purpose GPU pipeline. Each chip delivers 13.4 PFLOP/s of MXFP4 matrix compute against 216 GiB of HBM4 running at 15.4 TB/s, inside a 700 W package. Scaling is explicitly two-tier: a tightly coupled 128-ASIC domain at 600 GB/s, and a global 2,048-ASIC domain at 200 GB/s, which together form a 27 EFLOP/s, 432 TiB system.

Jalapeño is Gen 1 of a stated multi-generation platform. OpenAI's engineers took the design from initial RTL to tapeout in roughly nine months, and the company has described the next generations as targeting performance per watt first and then economical low-latency serving, with the goal of unlocking aggregate HBM bandwidth rather than simply adding raw bandwidth. Initial deployment is targeted for the end of 2026 starting through Microsoft Azure. QS Compute supplies LLM inference accelerators, GPU servers and rack-scale AI systems — request a quote.

Key Benefits

Inference-first silicon: MXFP4 matrix engines sized for LLM serving rather than training. Huge fast memory: 216 GiB HBM4 at 15.4 TB/s per chip. Rack-scale fabric: 600 GB/s local and 200 GB/s global domains across 2,048 chips. Rapid iteration: nine months from RTL to tapeout, with Gen 2 and Gen 3 already scoped.

Applications

Large-language-model inference at datacentre scale, long-context and retrieval-augmented serving, agentic AI workloads, chat and assistant platforms, and high-throughput token generation where cost per token dominates the deployment decision.

Request a Quote — OPENAI JALAPEÑO — BROADCOM-BUILT LLM INFERENCE ASIC

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

NVIDIA Rubin GPU — Next-Gen AI Accelerator AMD Instinct MI455X — 432GB HBM4 Accelerator NVIDIA Groq 3 LPU