Published: August 16, 2026 | Category: Buying Guide | QSCompute
Every AI inference server is ultimately a memory machine, not a compute machine. A model can only serve tokens from parameters and KV cache that actually reside in addressable memory — and for LLM serving, the KV cache is what explodes. A single H100 carries 80 GB of HBM3; a B200 pushes that to 192 GB of HBM3e. But when you need to hold a 70B-parameter model plus the KV cache for hundreds of concurrent sessions, the capacity wall arrives fast. This is where CXL (Compute Express Link) memory expansion and pooling enter the BOM in 2026 — adding terabytes of coherent DRAM beyond what a CPU's native DIMM slots allow.
Memory inside an inference node sits in a strict hierarchy, and each tier has a hard ceiling:
| Memory Tier | Typical Capacity | Latency | Cost per GB (Q3 2026) | Limitation |
|---|---|---|---|---|
| GPU HBM3e (H100/H200/B200) | 80–192 GB per GPU | ~1 µs | $15–$25 | Soldered, fixed, very expensive |
| Native DDR5 RDIMM | Up to ~3 TB per socket (12× 256 GB) | ~80–110 ns | $10–$14 | DIMM slot count caps capacity |
| CXL-attached DRAM | 2–4 TB per socket (8–12 expanders) | ~170–250 ns | $9–$12 | Higher latency, needs CXL-capable CPU |
| NVMe SSD (page cache) | 15–122 TB per drive | ~10–100 µs | $0.05–$0.10 | Too slow for live KV cache |
The pain point is specific: KV cache growth. As context windows stretch to 128K tokens and concurrency climbs, KV cache consumes memory far faster than model weights. A 70B model's weights are ~140 GB in FP16, but serving 500 concurrent sessions with 32K-token contexts can demand 2–4 TB of KV cache. Native DDR5 DIMM slots can't reach that economically, and HBM is the wrong tool — you don't want to pay HBM prices to park stale KV entries. CXL-attached DRAM slots in as the capacity tier in between.
CXL is a cache-coherent interconnect that rides the PCIe physical layer, exposing three device types: Type 1 (accelerators), Type 2 (accelerators with local memory), and Type 3 (memory expanders) — the DRAM-only devices this guide is about. The standard has moved quickly:
| Feature | CXL 2.0 | CXL 3.0 / 3.1 |
|---|---|---|
| Data rate | 32 GT/s (PCIe 5.0) | 64 GT/s (PCIe 6.0) |
| Memory expansion (one host) | Yes | Yes |
| Memory pooling via switch | Yes (single-level) | Yes (multi-level, fabric) |
| Memory sharing across hosts | No | Yes |
| Fabric-attached memory | No | Yes |
| Peer-to-peer (device-to-device) | Limited | Yes (CXL 3.1) |
| Typical expander capacity | 128–256 GB | 256 GB–2 TB |
| Maturity in 2026 | Shipping, battle-tested | Early, ecosystem forming |
The practical takeaway for buyers: CXL 2.0 memory expanders are real, shipping, and purchasable today. CXL 3.x brings pooling and fabric semantics that matter for multi-tenant datacenters, but the hardware ecosystem (switches, fabric managers) is still maturing in 2026. For a single-server capacity problem, CXL 2.0 is the pragmatic choice.
These terms are often conflated but solve different problems:
| Dimension | Expansion (Type 3 attached to one host) | Pooling (via CXL switch/fabric) |
|---|---|---|
| What it does | Adds DRAM capacity to one server | Shares a memory pool across multiple servers |
| Latency | ~170–250 ns (one hop) | ~250–400 ns (switch hop added) |
| Capacity flexibility | Fixed to the host | Elastic — reallocate at runtime |
| Complexity | Low — plug in an E3.S module | High — needs switch + fabric manager |
| Best for | Single inference node, KV cache, in-memory DB | Multi-tenant, bursty workloads, rack-scale tiering |
| Cost | Lower (module only) | Higher (switch + management) |
| Maturity 2026 | Production | Early adopter |
For the overwhelming majority of edge and enterprise AI inference deployments — single nodes or a handful of servers per site — expansion is the right first step. Pooling earns its complexity only when you have a rack-scale, multi-tenant memory problem.
| Product | Capacity | CXL Version | Form Factor | Controller | Street Price |
|---|---|---|---|---|---|
| Samsung CMM-D | 256 GB | CXL 2.0 | E3.S | Astera Labs Leo | ~$2,600–$3,200 |
| Micron CZ120 | 128 / 256 GB | CXL 2.0 | E3.S | Astera Labs Leo | ~$1,500 / $2,800 |
| SK hynix CMM-DDR5 | 256 GB | CXL 2.0 | E3.S | In-house + Astera | ~$2,800–$3,300 |
| Montage MXC-based modules | 128–256 GB | CXL 2.0 | E3.S | Montage MXC | ~$1,400–$2,600 |
| Astera Labs Leo (controller) | up to 2 TB | CXL 2.0/3.1 | (component) | — | $50–$90 |
Street pricing reflects the 2026 AI memory boom; all figures are allocation-dependent and subject to change.
The silicon underneath nearly all of these is the Astera Labs Leo Smart Memory Controller — the CXL memory controller that Samsung, Micron, and SK hynix build around. On the host side you need a CXL-capable CPU: AMD EPYC 9005 (Turin) and Intel Xeon 6 (Granite Rapids) both support CXL 2.0, as does AmpereOne on the Arm side.
CXL memory is not a drop-in replacement for native DDR5. The coherence protocol and PCIe hop add ~90–170 ns of round-trip latency over local NUMA memory. That makes it the wrong tier for a hot, latency-critical path — and exactly right for capacity-bound workloads:
| Workload | Best Memory Tier | Why |
|---|---|---|
| Active token generation (hot KV) | HBM / native DDR5 | Latency-sensitive per-token path |
| Stale / secondary KV cache | CXL-attached DRAM | Capacity-bound, rarely touched |
| Batch inference (offline) | CXL-attached DRAM | Throughput-bound, latency-tolerant |
| CPU-only LLM inference (70B+) | CXL-attached DRAM | Model simply doesn't fit in DIMM slots |
| In-memory embeddings / feature store | CXL-attached DRAM | Large, warm, not hot |
| On-prem fine-tuning checkpoints | NVMe | Cold, sequential |
The pattern is consistent: keep the hot working set in HBM or local DDR5, push the long tail into CXL memory. LLM serving frameworks in 2026 increasingly expose exactly this split — vLLM and SGLang can tier KV cache to a slower, larger pool rather than dropping context or spilling to disk.
| Question | Answer points to |
|---|---|
| Does the model + KV cache fit in current HBM + DDR5? | No change needed — don't over-buy |
| Are you capacity-bound on DRAM (not bandwidth)? | Add CXL 2.0 memory expanders |
| Is the workload latency-critical (real-time, sub-50 ms SLA)? | Stay on HBM / native DDR5 |
| CPU-side inference of a 70B+ model? | CXL expanders are the economical path |
| Multi-tenant, rack-scale, need elastic capacity? | Evaluate CXL 3.x pooling (early) |
| Budget trumps all, latency tolerant? | CXL expanders beat buying another GPU |
Rule of thumb: one CXL 256 GB expander (~$2,800) adds capacity at roughly $11/GB — cheaper than a second GPU's HBM and far cheaper than over-provisioning entire servers to get more DIMM slots. For a team running CPU-side LLM inference or KV-cache-heavy serving, CXL expansion is frequently the highest-ROI memory purchase available in 2026.
QSCompute stocks Samsung CMM-D, Micron CZ120, and SK hynix CXL memory expanders alongside CXL-ready AMD EPYC 9005 and Intel Xeon 6 servers, DDR5 RDIMM, and HBM3e — and our engineers can validate the exact CXL topology for your inference workload before you commit a BOM.
Need to break the DRAM capacity wall in your AI inference servers?
Our engineering team validates the exact CXL topology for your workload — KV-cache tiering, CPU-side inference, or in-memory feature stores — before you commit a BOM.
Contact: +86 137-1464-6179 | sherry@qscompute.com