Published: October 6, 2026 | Category: Technical Guide | QSCompute
The most common question a hardware vendor hears from a quantum team is the wrong one: "what GPU do I need for a quantum computer?" A quantum processor does not replace a GPU — it depends on one, and often on a rack of them. Two distinct classical workloads sit around every qubit array, and they pull the specification in opposite directions. The first is simulation: emulating a quantum circuit on classical hardware to develop algorithms, verify designs and calibrate a machine before it exists. The second is control and readout: generating the microwave and RF pulses that drive the qubits and classifying the faint signals that come back, inside a hard real-time deadline. One is throughput- and memory-bound; the other is latency-bound and barely touches a GPU. Buying one machine for both is the mistake this guide helps you avoid.
"Quantum simulation" is not one workload. State-vector simulation stores the full 2n-amplitude wavefunction and is exact but exponentially hungry. Tensor-network methods — matrix product states (MPS) and their two-dimensional cousins (PEPS) — exploit limited entanglement to reach far more qubits at the cost of approximation and structure-dependent performance. Quantum Monte Carlo and error-corrected logical-circuit simulation each have their own profile again. The right GPU depends entirely on which regime your team spends its hours in, so the first procurement question is not "H100 or B200?" but "state-vector or tensor-network?"
| Method | Qubits reachable | Precision profile | Binding constraint |
|---|---|---|---|
| Full state-vector | ~30–36 (VRAM-bound) | Complex128 dominant | Total VRAM, then GPU count |
| State-vector with QRAM / sparse | Adds a few qubits | Complex128 | Memory bandwidth |
| Tensor network (MPS) | Tens to hundreds (low entanglement) | Complex64 usually sufficient | Bond dimension, contraction order |
| Tensor network (PEPS / 2-D) | ~100+ on structured circuits | Complex64 | Contraction cost, interconnect |
| Quantum Monte Carlo | Hundreds of sites | FP64/FP32 mix | Aggregate sampling throughput |
| Error-corrected logical sim | Small logical, large physical | Mixed, heavy | VRAM + sustained FP64 |
A full state-vector of n qubits holds 2n complex amplitudes; in double precision that is 2n × 16 bytes. The wall arrives fast: 30 qubits needs about 16 GiB, 33 qubits about 128 GiB, and 36 qubits roughly a terabyte — which is precisely why big state-vector runs are sharded across a whole node of GPUs and stitched together over NVLink or InfiniBand. Simulators that offer a single-precision (complex64) mode double the qubit count a given card can hold, at the cost of numerical stability that some algorithms tolerate and some do not. Either way, the practical rule is that a state-vector machine is specified by aggregate VRAM, and its scaling ceiling is set by how well the shard-and-exchange step keeps up — the same interconnect problem that governs any tightly coupled HPC code.
Because complex128 state-vector math is FP64-heavy, data-centre cards with a strong double-precision ratio — the H100 and H200, at roughly 1:2 FP64:FP32 with a dedicated FP64 tensor path — remain the natural fit for large exact simulation, exactly as they are for computational chemistry and numerical weather prediction. Tensor-network work, by contrast, is usually happy in complex64 and behaves more like a dense linear-algebra or even machine-learning workload, which brings the cheaper high-FP32 cards into play. A programme that only does MPS calibration should not buy the same silicon as one running exact 34-qubit circuits. Settle the method, and the card class follows.
| Platform | VRAM | FP64 posture | Memory BW | Fit for quantum work |
|---|---|---|---|---|
| L40S | 48 GB ECC | Weak (~1:64) | 864 GB/s | Tensor-network (complex64), control-plane ML, visualisation |
| RTX 6000 Ada | 48 GB | Weak (~1:64) | 960 GB/s | Dev workstations, MPS/PEPS prototyping, pulse-shape ML |
| H100 SXM | 80 GB HBM3 | Strong (1:2) | 3,350 GB/s | Exact state-vector up to ~33 qubits, error-corrected sim |
| H200 SXM | 141 GB HBM3e | Strong (1:2) | 4,800 GB/s | Largest single-node state-vector, fewer shard exchanges |
| B200 SXM | 180 GB HBM3e | Strong (1:2) | 8,000 GB/s | Maximum qubit count per node, ensemble of circuit calibrations |
The readout and control path looks almost nothing like the simulation path. Pulse generation and qubit readout are closed-loop, deterministic workloads: a classical controller must synthesise an arbitrary-waveform pulse, measure the returned signal, classify the outcome and feed the result back into the next pulse — all inside a coherence-limited budget measured in hundreds of nanoseconds to a few microseconds. That job belongs to RFSoC and FPGA-based controllers with real-time converters, not to a data-centre GPU sitting behind a PCIe link and an OS scheduler. GPUs enter the control plane only where classification is genuinely neural — for example, a small network that discriminates qubit states from noisy readout traces, or that tunes pulse parameters — and even then it runs beside the deterministic controller, not in the feedback loop. Confusing these two planes is the most common architecture error in new quantum labs.
Any quantum programme that simulates at scale inherits the standard tightly coupled HPC problem. State-vector sharding exchanges amplitudes every gate, so the fabric between GPUs — NVLink and NVSwitch inside a node, NDR InfiniBand or RoCE between nodes — caps how far a run scales before communication dominates compute. Checkpoint I/O matters too: calibration sweeps and variational optimisations generate large volumes of intermediate results that stall the run if the filesystem cannot absorb the writes. And because the same cluster typically also runs the machine-learning models that process readout data and tune parameters, the node must serve both the FP64-heavy simulator and the tensor-shaped workloads on one platform — which is an argument for a mixed deployment rather than a uniform one.
| Dimension | Wrong choice | Recommended | Why it matters |
|---|---|---|---|
| Control & readout | GPU in the feedback loop | RFSoC/FPGA real-time controller + GPU alongside | Feedback deadline is microseconds; an OS-scheduled GPU cannot meet it |
| GPU–GPU fabric | PCIe Gen4 across sockets | NVLink / NVSwitch intra-node, NDR InfiniBand inter-node | Amplitude exchange every gate stalls a sharded state-vector |
| Precision | complex128 everywhere | complex64 for tensor networks, complex128 for exact sim | Halves VRAM and doubles qubit reach where algorithms allow |
| Checkpoint store | Single SATA SSD | PLP NVMe pool or parallel filesystem | Calibration sweeps stall the run without fast writes |
| Time & sync | Host clock only | IEEE 1588 PTP across controller and capture | Pulse timestamps only correlate with an accurate reference |
Building a classical simulation or control cluster for a quantum programme?
QSCompute assembles and burn-in tests GPU nodes for quantum simulation and control — H100, H200 and B200 SXM platforms with NVLink/NVSwitch, NDR InfiniBand and PLP NVMe storage, plus workstation-class L40S and RTX 6000 Ada nodes for tensor-network and readout-ML work. CUDA, cuQuantum and common quantum SDKs pre-installed. Volume pricing and DDP shipping worldwide.
Contact: +86 137-1464-6179 | info@qscompute.com