NVMe-oF & RDMA Storage Fabrics 2026 — NVMe/TCP vs RoCEv2 vs InfiniBand for Disaggregated AI Storage

Published: August 25, 2026 | Category: Technical | QSCompute

GPUs are idle more than anyone admits — and a large share of that idle time is spent waiting on 存储, not on compute. AI training jobs stream sharded datasets, write multi-terabyte checkpoints, and read them back after every preemption. When the storage sits in a separate enclosure from the GPU servers, the fabric connecting them — NVMe over Fabrics, or NVMe-oF — becomes the difference between a busy GPU and a starved one.

This guide explains the four NVMe-oF transports, what JBOF/EBOF disaggregation buys you, and how to choose a fabric for AI data pipelines without over-buying a lossless InfiniBand network you may not need.

What NVMe-oF Solves

Direct-attached NVMe SSDs deliver sub-20 µs latency inside a single server, but that storage dies with the node and cannot be shared. NVMe-oF extends the NVMe protocol across a network, so many servers can share a pool of flash with most of the local-NVMe performance. The trade-offs live entirely in which transport you choose.

TransportNetworkTypical LatencySpecial HardwareBest For
NVMe/TCPStandard Ethernet (25–100 GbE)150–250 µsNone (software initiator)Cost-sensitive, general-purpose AI storage
NVMe/RDMA (RoCEv2)Ethernet + RoCEv2, lossless (PFC/ECN)50–150 µsRDMA-capable NICs (ConnectX, E810)High-throughput training data pipelines
NVMe/RDMA (InfiniBand)InfiniBand HDR/NDR (200/400 Gb/s)1–10 µsInfiniBand HCAs + switchesLatency-critical checkpointing, HPC-style AI
NVMe/FCFibre Channel (32/64 GFC)100–200 µsFC HBAs + SAN switchesEnterprise SAN shops consolidating onto NVMe

JBOF and EBOF: The Disaggregated Building Block

JBOF (Just a Bunch Of Flash) is a headless enclosure full of NVMe drives with no compute — it exposes raw drives to the fabric. EBOF (Ethernet Bunch Of Flash) does the same over standard Ethernet, usually with an onboard controller that presents the drives as NVMe-oF targets. Because there is no x86 storage controller in the data path, EBOF appliances (from vendors like Supermicro, Western Digital OpenFlex, and Lightbits) can saturate 100 GbE with a handful of drives and add capacity independently of compute.

Why it matters for AI: a single EBOF shelf can feed dozens of GPU servers simultaneously, and capacity scales by adding shelves instead of buying bigger servers. Checkpoints land on shared flash instead of filling a single node's local disks.

Picking a Fabric for AI Workloads

NVMe/TCP first. For most AI training and inference pipelines, NVMe/TCP on 25–100 GbE is the pragmatic default. It needs no RDMA NICs, no lossless-network tuning, and no separate management domain — you get shared NVMe storage over the Ethernet you already run. Modern software initiators and offload-capable NICs (NVMe/TCP offload on Mellanox ConnectX and Intel E810) close much of the latency gap.

RoCEv2 when throughput dominates. If your dataset streaming or checkpoint workload is bandwidth-bound and you already run a RoCEv2 fabric for the GPU cluster, NVMe/RDMA cuts CPU overhead and latency. The cost is operational: RoCEv2 needs Priority Flow Control or ECN configured correctly, and a misconfigured lossless network causes packet storms that are painful to debug.

InfiniBand when latency is the product. Frontier labs that run InfiniBand NDR for GPU-to-GPU traffic extend the same fabric to storage, getting single-digit-microsecond NVMe access. It is the highest-performance option and by far the most expensive — only justified when storage latency is a measurable training bottleneck, not by default.

FC for the SAN incumbents. If your organization has a Fibre Channel SAN team, NVMe/FC lets you modernize to NVMe without abandoning the fabric, tools, and zoning model you already trust.

Realistic Costs and Sizing

An EBOF enclosure of 12–24 U.2/E3.S NVMe drives typically runs $15,000–40,000 before drives; NVMe/TCP needs no extra NICs, while RoCEv2 adds RDMA NICs (~$800–1,500 per 100 GbE port) and InfiniBand NDR switch ports run several thousand dollars each. For a mid-size training cluster, starting on NVMe/TCP and upgrading the storage fabric to RDMA only when profiling proves the storage is the constraint is the lowest-regret path.

Planning a shared NVMe-oF storage pool for your AI cluster?

QSCompute supplies EBOF/JBOF enclosures, NVMe-oF targets, RDMA NICs, and InfiniBand/Ethernet switches — pre-integrated, cabled, and validated with your GPU servers so the storage fabric is never the bottleneck.

Contact: +86 137-1464-6179 | sherry@qscompute.com