Specifications

Generation

HBM4, with HBM4E extending the roadmap using customizable base logic dies

Bus interface

2048-pin interface — twice the bus width of HBM3E

Pin speed

Greater than 11.0 Gb/s per pin, against 6.4 GT/s on the HBM3E generation

Bandwidth

Greater than 2.8 TB/s per stack, more than double the previous generation

12-high capacity

36 GB per stack, matching HBM3E 12-high density

16-high capacity

48 GB per stack, sampling to customers in 2026

Bandwidth gain

Roughly 2.3x the bandwidth of Micron HBM3E 12-high

Power efficiency

More than 20% better pJ/bit versus HBM3E under comparable conditions

DRAM process

1-beta for volume ramp with the 1-gamma node, Micron's first EUV DRAM, for HBM4E

Base die

In-house CMOS base die optimised for power and signal integrity, plus a TSMC co-developed customizable logic die for HBM4E

Stack construction

DRAM dies with through-silicon vias stacked on a logic base die, capped by a thicker top die

Capacity/bandwidth split

Capacity sets how much of a problem fits in memory; bandwidth sets how fast it can be solved

Interfaces

Supports PCIe Gen6 SSDs and SOCAMM2 modules alongside GPU and xPU HBM interfaces

Design wins

In high-volume production for NVIDIA Vera Rubin generation AI accelerators

Memory hierarchy role

Complements LPDDR5 and DDR5 system memory rather than replacing it

Workloads

Long-context reasoning, multimodal AI, agentic multi-agent systems and HPC simulation

Overview

Micron's HBM4 widens the high-bandwidth memory interface to 2048 bits — double the HBM3E bus — and pushes pin speeds beyond 11.0 Gb/s to deliver more than 2.8 TB/s per stack. That is roughly 2.3x the bandwidth of Micron's HBM3E 12-high parts at similar or better pJ/bit efficiency, and it is what allows inference servers to stream the terabyte-scale context windows that long-context reasoning models demand.

The 12-high configuration carries 36 GB per stack with the same density as the previous generation, while a 16-high 48 GB stack entered customer sampling in 2026. Micron builds the stack from through-silicon-via DRAM dies on a logic base die, using 1-beta DRAM for the volume ramp and its 1-gamma EUV node for the next step.

HBM4E is the follow-on generation, and its defining change is customisation: Micron ships an in-house CMOS base die with HBM4E and has co-developed customizable base logic dies with TSMC, letting accelerator designers put their own compute, test or interface logic into the memory stack. Micron states HBM4 is in high-volume production for the NVIDIA Vera Rubin generation and is also designed into PCIe Gen6 SSDs and SOCAMM2 modules.

Key Benefits

Greater than 2.8 TB/s of bandwidth per stack on a 2048-bit interface removes the memory wall for long-context inference, multimodal models and multi-agent serving. 2.3x the bandwidth of HBM3E with over 20% better pJ/bit improves tokens per second per watt at rack level. In-house CMOS base die plus TSMC co-developed customizable HBM4E logic lets accelerator teams embed differentiating logic in the memory stack, while 36 GB 12-high now and 48 GB 16-high sampling in 2026 give a clear capacity roadmap.

Applications

AI GPU and xPU main memory in training and inference accelerators; long-context reasoning and agentic AI servers; multimodal models processing text, image, video and sensor data; HPC and scientific simulation; PCIe Gen6 SSD caching; and SOCAMM2-based AI server memory hierarchies.

Request a Quote — MICRON HBM4E — CUSTOMIZABLE BASE DIE HBM FOR AI GPUS

QS Compute — global B2B supply of AI computing hardware, edge AI systems and accelerators. Volume pricing, 15-day sample lead time.

Get Your Quote →

Related Products

Micron HBM4 16-High 48GB Micron HBM3E 12-High 36GB Samsung HBM4E SK hynix HBM4E 48GB Micron HBM4 48GB