Specifications
Generation
HBM4, with HBM4E extending the roadmap using customizable base logic dies
Bus interface
2048-pin interface — twice the bus width of HBM3E
Pin speed
Greater than 11.0 Gb/s per pin, against 6.4 GT/s on the HBM3E generation
Bandwidth
Greater than 2.8 TB/s per stack, more than double the previous generation
12-high capacity
36 GB per stack, matching HBM3E 12-high density
16-high capacity
48 GB per stack, sampling to customers in 2026
Bandwidth gain
Roughly 2.3x the bandwidth of Micron HBM3E 12-high
Power efficiency
More than 20% better pJ/bit versus HBM3E under comparable conditions
DRAM process
1-beta for volume ramp with the 1-gamma node, Micron's first EUV DRAM, for HBM4E
Base die
In-house CMOS base die optimised for power and signal integrity, plus a TSMC co-developed customizable logic die for HBM4E
Stack construction
DRAM dies with through-silicon vias stacked on a logic base die, capped by a thicker top die
Capacity/bandwidth split
Capacity sets how much of a problem fits in memory; bandwidth sets how fast it can be solved
Interfaces
Supports PCIe Gen6 SSDs and SOCAMM2 modules alongside GPU and xPU HBM interfaces
Design wins
In high-volume production for NVIDIA Vera Rubin generation AI accelerators
Memory hierarchy role
Complements LPDDR5 and DDR5 system memory rather than replacing it
Workloads
Long-context reasoning, multimodal AI, agentic multi-agent systems and HPC simulation
Overview
Micron's HBM4 widens the high-bandwidth memory interface to 2048 bits — double the HBM3E bus — and pushes pin speeds beyond 11.0 Gb/s to deliver more than 2.8 TB/s per stack. That is roughly 2.3x the bandwidth of Micron's HBM3E 12-high parts at similar or better pJ/bit efficiency, and it is what allows inference servers to stream the terabyte-scale context windows that long-context reasoning models demand.
The 12-high configuration carries 36 GB per stack with the same density as the previous generation, while a 16-high 48 GB stack entered customer sampling in 2026. Micron builds the stack from through-silicon-via DRAM dies on a logic base die, using 1-beta DRAM for the volume ramp and its 1-gamma EUV node for the next step.
HBM4E is the follow-on generation, and its defining change is customisation: Micron ships an in-house CMOS base die with HBM4E and has co-developed customizable base logic dies with TSMC, letting accelerator designers put their own compute, test or interface logic into the memory stack. Micron states HBM4 is in high-volume production for the NVIDIA Vera Rubin generation and is also designed into PCIe Gen6 SSDs and SOCAMM2 modules.
Key Benefits
Greater than 2.8 TB/s of bandwidth per stack on a 2048-bit interface removes the memory wall for long-context inference, multimodal models and multi-agent serving. 2.3x the bandwidth of HBM3E with over 20% better pJ/bit improves tokens per second per watt at rack level. In-house CMOS base die plus TSMC co-developed customizable HBM4E logic lets accelerator teams embed differentiating logic in the memory stack, while 36 GB 12-high now and 48 GB 16-high sampling in 2026 give a clear capacity roadmap.
Applications
AI GPU and xPU main memory in training and inference accelerators; long-context reasoning and agentic AI servers; multimodal models processing text, image, video and sensor data; HPC and scientific simulation; PCIe Gen6 SSD caching; and SOCAMM2-based AI server memory hierarchies.
Request a Quote — MICRON HBM4E — CUSTOMIZABLE BASE DIE HBM FOR AI GPUS
QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.
Get Your Quote →