Published: August 25, 2026 | Category: Technical | QSCompute
GPUs are idle more than anyone admits — and a large share of that idle time is spent waiting on 存储, not on compute. AI training jobs stream sharded datasets, write multi-terabyte checkpoints, and read them back after every preemption. When the storage sits in a separate enclosure from the GPU servers, the fabric connecting them — NVMe over Fabrics, or NVMe-oF — becomes the difference between a busy GPU and a starved one.
This guide explains the four NVMe-oF transports, what JBOF/EBOF disaggregation buys you, and how to choose a fabric for AI data pipelines without over-buying a lossless InfiniBand network you may not need.
Direct-attached NVMe SSDs deliver sub-20 µs latency inside a single server, but that storage dies with the node and cannot be shared. NVMe-oF extends the NVMe protocol across a network, so many servers can share a pool of flash with most of the local-NVMe performance. The trade-offs live entirely in which transport you choose.
| Transport | Network | Typical Latency | Special Hardware | Best For |
|---|---|---|---|---|
| NVMe/TCP | Standard Ethernet (25–100 GbE) | 150–250 µs | None (software initiator) | Cost-sensitive, general-purpose AI storage |
| NVMe/RDMA (RoCEv2) | Ethernet + RoCEv2, lossless (PFC/ECN) | 50–150 µs | RDMA-capable NICs (ConnectX, E810) | High-throughput training data pipelines |
| NVMe/RDMA (InfiniBand) | InfiniBand HDR/NDR (200/400 Gb/s) | 1–10 µs | InfiniBand HCAs + switches | Latency-critical checkpointing, HPC-style AI |
| NVMe/FC | Fibre Channel (32/64 GFC) | 100–200 µs | FC HBAs + SAN switches | Enterprise SAN shops consolidating onto NVMe |
JBOF (Just a Bunch Of Flash) is a headless enclosure full of NVMe drives with no compute — it exposes raw drives to the fabric. EBOF (Ethernet Bunch Of Flash) does the same over standard Ethernet, usually with an onboard controller that presents the drives as NVMe-oF targets. Because there is no x86 storage controller in the data path, EBOF appliances (from vendors like Supermicro, Western Digital OpenFlex, and Lightbits) can saturate 100 GbE with a handful of drives and add capacity independently of compute.
NVMe/TCP first. For most AI training and inference pipelines, NVMe/TCP on 25–100 GbE is the pragmatic default. It needs no RDMA NICs, no lossless-network tuning, and no separate management domain — you get shared NVMe storage over the Ethernet you already run. Modern software initiators and offload-capable NICs (NVMe/TCP offload on Mellanox ConnectX and Intel E810) close much of the latency gap.
RoCEv2 when throughput dominates. If your dataset streaming or checkpoint workload is bandwidth-bound and you already run a RoCEv2 fabric for the GPU cluster, NVMe/RDMA cuts CPU overhead and latency. The cost is operational: RoCEv2 needs Priority Flow Control or ECN configured correctly, and a misconfigured lossless network causes packet storms that are painful to debug.
InfiniBand when latency is the product. Frontier labs that run InfiniBand NDR for GPU-to-GPU traffic extend the same fabric to storage, getting single-digit-microsecond NVMe access. It is the highest-performance option and by far the most expensive — only justified when storage latency is a measurable training bottleneck, not by default.
FC for the SAN incumbents. If your organization has a Fibre Channel SAN team, NVMe/FC lets you modernize to NVMe without abandoning the fabric, tools, and zoning model you already trust.
An EBOF enclosure of 12–24 U.2/E3.S NVMe drives typically runs $15,000–40,000 before drives; NVMe/TCP needs no extra NICs, while RoCEv2 adds RDMA NICs (~$800–1,500 per 100 GbE port) and InfiniBand NDR switch ports run several thousand dollars each. For a mid-size training cluster, starting on NVMe/TCP and upgrading the storage fabric to RDMA only when profiling proves the storage is the constraint is the lowest-regret path.
Planning a shared NVMe-oF storage pool for your AI cluster?
QSCompute supplies EBOF/JBOF enclosures, NVMe-oF targets, RDMA NICs, and InfiniBand/Ethernet switches — pre-integrated, cabled, and validated with your GPU servers so the storage fabric is never the bottleneck.
Contact: +86 137-1464-6179 | sherry@qscompute.com