GPU Hardware for Quantum Computing Simulation & Control 2026 — State-Vector, Tensor-Network & Cryogenic Control

Published: October 6, 2026 | Category: Technical Guide | QSCompute

The most common question a hardware vendor hears from a quantum team is the wrong one: "what GPU do I need for a quantum computer?" A quantum processor does not replace a GPU — it depends on one, and often on a rack of them. Two distinct classical workloads sit around every qubit array, and they pull the specification in opposite directions. The first is simulation: emulating a quantum circuit on classical hardware to develop algorithms, verify designs and calibrate a machine before it exists. The second is control and readout: generating the microwave and RF pulses that drive the qubits and classifying the faint signals that come back, inside a hard real-time deadline. One is throughput- and memory-bound; the other is latency-bound and barely touches a GPU. Buying one machine for both is the mistake this guide helps you avoid.

Simulation Regimes and Their Hardware Appetite

"Quantum simulation" is not one workload. State-vector simulation stores the full 2n-amplitude wavefunction and is exact but exponentially hungry. Tensor-network methods — matrix product states (MPS) and their two-dimensional cousins (PEPS) — exploit limited entanglement to reach far more qubits at the cost of approximation and structure-dependent performance. Quantum Monte Carlo and error-corrected logical-circuit simulation each have their own profile again. The right GPU depends entirely on which regime your team spends its hours in, so the first procurement question is not "H100 or B200?" but "state-vector or tensor-network?"

MethodQubits reachablePrecision profileBinding constraint
Full state-vector~30–36 (VRAM-bound)Complex128 dominantTotal VRAM, then GPU count
State-vector with QRAM / sparseAdds a few qubitsComplex128Memory bandwidth
Tensor network (MPS)Tens to hundreds (low entanglement)Complex64 usually sufficientBond dimension, contraction order
Tensor network (PEPS / 2-D)~100+ on structured circuitsComplex64Contraction cost, interconnect
Quantum Monte CarloHundreds of sitesFP64/FP32 mixAggregate sampling throughput
Error-corrected logical simSmall logical, large physicalMixed, heavyVRAM + sustained FP64

The VRAM Wall Is Arithmetic, Not Marketing

A full state-vector of n qubits holds 2n complex amplitudes; in double precision that is 2n × 16 bytes. The wall arrives fast: 30 qubits needs about 16 GiB, 33 qubits about 128 GiB, and 36 qubits roughly a terabyte — which is precisely why big state-vector runs are sharded across a whole node of GPUs and stitched together over NVLink or InfiniBand. Simulators that offer a single-precision (complex64) mode double the qubit count a given card can hold, at the cost of numerical stability that some algorithms tolerate and some do not. Either way, the practical rule is that a state-vector machine is specified by aggregate VRAM, and its scaling ceiling is set by how well the shard-and-exchange step keeps up — the same interconnect problem that governs any tightly coupled HPC code.

The FP64 Question for Quantum

Because complex128 state-vector math is FP64-heavy, data-centre cards with a strong double-precision ratio — the H100 and H200, at roughly 1:2 FP64:FP32 with a dedicated FP64 tensor path — remain the natural fit for large exact simulation, exactly as they are for computational chemistry and numerical weather prediction. Tensor-network work, by contrast, is usually happy in complex64 and behaves more like a dense linear-algebra or even machine-learning workload, which brings the cheaper high-FP32 cards into play. A programme that only does MPS calibration should not buy the same silicon as one running exact 34-qubit circuits. Settle the method, and the card class follows.

PlatformVRAMFP64 postureMemory BWFit for quantum work
L40S48 GB ECCWeak (~1:64)864 GB/sTensor-network (complex64), control-plane ML, visualisation
RTX 6000 Ada48 GBWeak (~1:64)960 GB/sDev workstations, MPS/PEPS prototyping, pulse-shape ML
H100 SXM80 GB HBM3Strong (1:2)3,350 GB/sExact state-vector up to ~33 qubits, error-corrected sim
H200 SXM141 GB HBM3eStrong (1:2)4,800 GB/sLargest single-node state-vector, fewer shard exchanges
B200 SXM180 GB HBM3eStrong (1:2)8,000 GB/sMaximum qubit count per node, ensemble of circuit calibrations

The Control Plane Is a Latency Problem, Not a Throughput One

The readout and control path looks almost nothing like the simulation path. Pulse generation and qubit readout are closed-loop, deterministic workloads: a classical controller must synthesise an arbitrary-waveform pulse, measure the returned signal, classify the outcome and feed the result back into the next pulse — all inside a coherence-limited budget measured in hundreds of nanoseconds to a few microseconds. That job belongs to RFSoC and FPGA-based controllers with real-time converters, not to a data-centre GPU sitting behind a PCIe link and an OS scheduler. GPUs enter the control plane only where classification is genuinely neural — for example, a small network that discriminates qubit states from noisy readout traces, or that tunes pulse parameters — and even then it runs beside the deterministic controller, not in the feedback loop. Confusing these two planes is the most common architecture error in new quantum labs.

Interconnect, Storage and the Rest of the Node

Any quantum programme that simulates at scale inherits the standard tightly coupled HPC problem. State-vector sharding exchanges amplitudes every gate, so the fabric between GPUs — NVLink and NVSwitch inside a node, NDR InfiniBand or RoCE between nodes — caps how far a run scales before communication dominates compute. Checkpoint I/O matters too: calibration sweeps and variational optimisations generate large volumes of intermediate results that stall the run if the filesystem cannot absorb the writes. And because the same cluster typically also runs the machine-learning models that process readout data and tune parameters, the node must serve both the FP64-heavy simulator and the tensor-shaped workloads on one platform — which is an argument for a mixed deployment rather than a uniform one.

DimensionWrong choiceRecommendedWhy it matters
Control & readoutGPU in the feedback loopRFSoC/FPGA real-time controller + GPU alongsideFeedback deadline is microseconds; an OS-scheduled GPU cannot meet it
GPU–GPU fabricPCIe Gen4 across socketsNVLink / NVSwitch intra-node, NDR InfiniBand inter-nodeAmplitude exchange every gate stalls a sharded state-vector
Precisioncomplex128 everywherecomplex64 for tensor networks, complex128 for exact simHalves VRAM and doubles qubit reach where algorithms allow
Checkpoint storeSingle SATA SSDPLP NVMe pool or parallel filesystemCalibration sweeps stall the run without fast writes
Time & syncHost clock onlyIEEE 1588 PTP across controller and capturePulse timestamps only correlate with an accurate reference

Selection Rules

Building a classical simulation or control cluster for a quantum programme?

QSCompute assembles and burn-in tests GPU nodes for quantum simulation and control — H100, H200 and B200 SXM platforms with NVLink/NVSwitch, NDR InfiniBand and PLP NVMe storage, plus workstation-class L40S and RTX 6000 Ada nodes for tensor-network and readout-ML work. CUDA, cuQuantum and common quantum SDKs pre-installed. Volume pricing and DDP shipping worldwide.

Contact: +86 137-1464-6179 | info@qscompute.com