Published: September 24, 2026 | Category: Technical Guide | QSCompute
Two buyers arrive at the same GPU decision from opposite directions. One is deploying inference and cares about TOPS per watt in a sealed enclosure. The other is a chemistry or drug-discovery group that wants a simulation to finish before the grant runs out, and cares about none of that. Sequential throughput of a numerical integrator, the precision that integrator runs in, and the capacity of a system that has to fit entirely in VRAM decide whether the second buyer's money worked.
This guide is about that second buyer: the workstation or small on-premise cluster running molecular dynamics, docking campaigns, free-energy calculations and increasingly machine-learned potentials. The general shape is not the same as an inference box, and the sizing errors are different too.
An MD engine spends nearly all its time in one loop: compute forces between atoms, integrate the equations of motion, repeat millions of times. The expensive part is the force calculation and the long-range electrostatics handled by particle-mesh Ewald methods, which is FFT-heavy but not tensor-shaped. There is no batch of independent tokens to fill a matrix unit; there is a single large system stepped serially, and the timestep is bounded by physics rather than by hardware.
That produces three hardware questions, in this order: does the simulation fit in the card's memory, at what precision can the force field be evaluated, and how fast can the card read and write memory. Peak AI throughput is close to irrelevant.
| Workload | Dominant hardware constraint | Sizing note |
|---|---|---|
| MD integration (GROMACS, AMBER, OpenMM, NAMD, LAMMPS) | Memory bandwidth and FP32/FP64 throughput; VRAM capacity for the whole system plus neighbour lists | One simulation per GPU; buy capacity before clock speed |
| Long-range electrostatics (PME) | FFT throughput, host–device transfer if the mesh is offloaded | Amortised inside the same timestep — counts as bandwidth pressure |
| Enhanced sampling & free-energy campaigns (replica exchange, FEP, metadynamics) | GPU count, not GPU speed — hundreds of parallel windows or replicas | Many mid-range cards beat one flagship for the same budget |
| Docking & virtual screening | Throughput on many small independent jobs; occupancy matters more than bandwidth | Embarrassingly parallel; good fit for a scheduling layer over commodity cards |
| Neural network potentials / MLIP (ANI, MACE, NequIP) | Genuinely tensor-shaped: FP16/BF16 matrix throughput and VRAM | The one chemistry workload that looks like AI inference |
| QM/MM and electronic structure | High-precision math; often CPU- or hybrid-bound | Do not assume a GPU-heavy configuration is optimal |
| Trajectory analysis & visualisation | I/O and memory, occasionally GPU-accelerated analysis kernels | Almost always limited by storage, not compute |
Start from atoms, because everything else follows. A modest solvated protein system of 100,000 atoms costs 100,000 × 3 coordinates × 4 bytes to write one frame — about 1.2 MB. Write that frame every 10 ps over a 1 µs trajectory and the run produces 100,000 frames, on the order of 120 GB, for a single production run. Write every 100 ps instead and the same science costs 12 GB. That single choice about sampling interval moves the storage bill by a factor of ten and is usually made by accident.
Precision is the second gate. Most MD engines run a "mixed" scheme: forces in single precision with double-precision accumulation of the critical terms, which is why a card's FP64 rate still matters even though the bulk of the arithmetic is FP32. On many GeForce-class parts FP64 executes at a small fraction of FP32 rate — ratios as coarse as 1:32 or 1:64 appear across generations — so a card advertised on inference throughput can be a disappointing integrator. If the method genuinely needs FP64 (certain free-energy estimators, polarisable force fields, some QM/MM coupling), check the FP64 column and nothing else.
The third gate is VRAM, and it fails quietly. When a system plus its neighbour lists and scratch buffers exceed card memory, the engine either refuses to run or spills to host memory, at which point throughput collapses by an order of magnitude with no error message printed. Capacity is therefore a hard floor: choose the card that fits the largest system the group intends to study, not the fastest card that fits today's smallest one.
| Tier | Shape | What it suits | Typical street price band |
|---|---|---|---|
| Single-workstation card | One 48–96 GB card in a tower, 8–16 host cores, 128–256 GB ECC RAM | One large system at a time; visualisation and analysis interspersed | $5,000–15,000 |
| Dual-GPU tower | Two cards, 1–2 kW supply, PCIe lane budget planned deliberately | One large run plus one campaign, or two scientists sharing a bench instrument | $12,000–35,000 |
| Small GPU server | 4 GPUs with NVLink or a fast fabric, redundant PSU, rack or tower | Replica exchange and FEP campaigns; the department's shared resource | $40,000–150,000 |
| Entry / teaching card | Single 16–24 GB card, modest supply | Small systems, docking, teaching, MLIP inference | $1,500–4,000 |
| Burst capacity | Cloud or an institutional HPC allocation | Campaigns that must finish in a week rather than a quarter | Per-GPU-hour, no capital |
A chemistry workstation is not a GPU with a computer attached. Host RAM should hold at least twice the largest system so that analysis, trajectory loading and the engine's own buffers do not evict each other; ECC matters when a run is a week long and a silent bit flip invalidates a month of campaign work. PCIe lane allocation has to be planned before purchase: two cards on a consumer chipset may each receive half the lanes they want, and the loss shows up as a stall between timesteps rather than as a benchmark difference.
Storage is the part chemistry groups undersize most often, because trajectories are written continuously for the entire duration of a run and cannot be compressed on the fly without cost.
| Data class | Volume to plan for | Design requirement |
|---|---|---|
| Live trajectory output | ~1.2 MB per 100,000-atom frame; tens to hundreds of GB per run | Local NVMe with a sustained-write claim, not a burst number |
| Checkpoint / restart files | Full system state every 5–15 minutes; similar size to the system image | Low latency, power-loss protected — a lost checkpoint is a lost week |
| Trajectory archive and retention | Multi-TB per project over a grant period | Tiered storage with integrity checking; hashes for reproducibility |
| Analysis scratch | 2–5× the trajectory being analysed | Fast local scratch, rebuilt freely; never the archive |
| Virtual screening results | Small (scores and poses), but millions of small files | Metadata handling, not bandwidth |
| Container images and engine builds | Hundreds of GB, versioned and pinned | Reproducibility requirement: pin the exact engine and CUDA build per publication |
QSCompute supplies the workstation and server layer for computational chemistry — single- and multi-GPU towers and 4U servers, RTX PRO 6000 Blackwell 96 GB and L40S-class cards, industrial NVMe scratch with power-loss protection, ECC DRAM, and the PCIe lane budgeting that keeps two cards from halving each other. We size against your largest system and sampling interval, not against a synthetic benchmark. DDP shipping to 85+ countries.
Sizing a chemistry GPU and not sure the system will fit?
Send us the largest system you intend to simulate, the engine and precision mode, and your sampling interval — our engineers return a card, memory and storage specification with the failure modes called out.
Contact: +86 137-1464-6179 | info@qscompute.com