Published: September 2, 2026 | Category: Buying Guide | QSCompute
A single NFS server moving data at 1.1 GB/s cannot feed a multi-GPU training cluster. A published 32-GPU benchmark measured checkpoint writes taking 127 seconds on NFS — with the cluster idle the whole time — versus 3–5 seconds on a parallel filesystem. Parallel filesystems fix this by striping files across many servers, so bandwidth scales with the number of storage nodes instead of stopping at one. This guide compares the four serious options for AI training in 2026 — Lustre, IBM Storage Scale (GPFS), Weka (NeuralMesh), and BeeGFS — with real throughput numbers, licensing models, and a sizing method you can use before talking to a vendor.
Training is storage-hungry in two directions. Data loading pulls the dataset through the cluster: an 8-GPU H100 node can demand 30–50 GB/s when dataloaders run hot, and a 32-GPU cluster wants hundreds of GB/s aggregated. Checkpointing is the second, sharper problem — every N steps each rank writes its weights, optimizer state, and RNG state, and the cluster stalls until the write finishes. A single 10 GbE NFS server sustains roughly 1.1 GB/s; a 25 GbE server with one NVMe drive tops out near 2–4 GB/s. Either way, a 400 GB checkpoint takes minutes — and every minute is a full cluster sitting idle.
| Storage setup (32-GPU cluster) | Aggregated write | Checkpoint time | GPU idle during checkpoint |
|---|---|---|---|
| Single NFS server, 10 GbE | ~1.1 GB/s | ~127 s | ~21% |
| BeeOND (32× local NVMe) | ~28 GB/s | ~5 s | ~0.8% |
| Lustre (32 OSTs, RoCEv2) | ~38 GB/s | ~3.7 s | ~0.6% |
| Weka (32 GPUs, distributed) | ~44 GB/s | ~3.2 s | ~0.5% |
Two structural facts push buyers toward parallel POSIX filesystems rather than object storage. First, the training stack — PyTorch DataLoader, checkpoint libraries, Horovod, most fine-tuning code — assumes POSIX paths; MinIO and Ceph require gateway shims that add latency. Second, a parallel filesystem's bandwidth scales linearly: add storage nodes, add throughput. The comparison table below shows how the four options differ.
| Lustre | IBM Storage Scale (GPFS) | Weka (NeuralMesh) | BeeGFS | |
|---|---|---|---|---|
| License | Open source (GPL); paid support (Whamcloud/DDN) | Commercial, per usable TiB (Data Access Edition) | Commercial, per usable TB subscription (1/3/5 yr) | Open source (GPL); paid support (DDN) |
| Metadata architecture | MDS/MDT pairs (DNE for striped MDTs) | Distributed, enterprise-hardened | Fully distributed, per-core metadata | Horizontal metadata servers |
| Typical scale | 100s of nodes, TB/s class | 10s–100s of nodes | 10s of nodes, GB/s per node | 10–100+ nodes, easy to grow |
| Strengths | HPC standard, proven at extreme scale | Policy tiering, small-file IOPS, enterprise data management | NVMe-native, ~30 GB/s and >1M IOPS per node, GPU Direct Storage, container-native CSI | Easiest to deploy and operate, low entry cost, strong AI track record |
| Watch out for | Ops expertise required; metadata bottleneck at scale | License cost on the full usable capacity | Subscription cost compounds on large clusters | Smaller vendor ecosystem than Lustre at extreme scale |
Lustre remains the default for HPC-scale AI (it runs most Top500 supercomputers) and its open-source license keeps the license line at zero — but it demands storage expertise, and its metadata servers are the classic bottleneck. IBM Storage Scale is the enterprise pick: capacity-based licensing counted per usable TiB (clients are free), policy-based tiering to object storage, and the strongest small-file performance of the four; expect the highest license line item. Weka (rebranded NeuralMesh in 2026) is software-only NVMe storage — reference designs from Supermicro and HPE show ~30 GB/s and over 1M IOPS per node at sub-200µs latency, with CSI and GPU Direct Storage built in; you pay a per-TB subscription on top of the hardware. BeeGFS is the pragmatic open-source choice: metadata servers scale horizontally with no single-active-MDS constraint, deployment takes a day, and DDN backs it commercially — which is why it shows up in a growing number of production AI clusters.
Size for checkpoint write budget first, then data-load read bandwidth. A practical rule: the filesystem should write a full checkpoint in under 60 seconds. A 400 GB checkpoint therefore needs roughly 7 GB/s of sustained write; a 1 TB checkpoint needs 17+ GB/s. Then add read headroom — most teams plan for 2–3× the checkpoint write figure as sustained read bandwidth for dataloaders.
Hardware math is straightforward. A single NVMe storage node with eight Gen5 drives delivers roughly 10–14 GB/s of sequential bandwidth and ~60 TB raw capacity; eight such nodes give you ~100 GB/s aggregated. Metadata needs two nodes (an HA pair) on Lustre and BeeGFS; Weka spreads metadata across all nodes. The fabric matters: RoCEv2 at 100/200/400 GbE or InfiniBand NDR, with NVMe-oF when the drives sit in separate JBODs — the network layer is covered in our NVMe-oF and cluster-networking guides.
Budget planning for Q3 2026: NVMe storage nodes run roughly $25–45k each (EPYC or Xeon, 8× Gen4/5 NVMe, dual 100/200 GbE NICs), landing at about $150–300 per usable TB of hardware. On top of that: Weka charges a per-TB subscription (1/3/5-year terms), IBM Scale licenses per usable TiB, and the open-source pair (Lustre, BeeGFS) charges only for support — typically a per-node annual contract. As a managed benchmark, AWS FSx for Lustre lists at $0.17–$0.30/GB/month depending on throughput tier.
The pattern across all four: the filesystem is now a first-class part of the GPU cluster budget, not an afterthought. Size it with the checkpoint math above, and you will never stall a cluster waiting on storage again.
Need a training cluster with storage that keeps up?
QSCompute builds GPU clusters and their parallel storage: EPYC/Xeon servers, Gen4/Gen5 NVMe storage nodes, 100/200/400 GbE RDMA fabrics, and burn-in validated at your checkpoint target. Tell us your GPU count and checkpoint size — we will size the filesystem that never idles your GPUs.
Contact: +86 137-1464-6179 | info@qscompute.com