Published: September 20, 2026 | Category: Buying Guide | QSCompute
A laboratory rarely sets out to buy storage. It buys a sequencer, a mass spectrometer or a slide scanner, and three years later discovers that the purchase came with a permanently growing data estate and a regulatory obligation to keep the raw files alive for a decade or more. Unlike a factory, where telemetry can be aggregated and discarded, a regulated lab has to prove that the original bytes are intact, attributable and retrievable. This guide covers the real volumes per bench, the retention floors by regime, and the storage architecture that satisfies both.
Laboratory storage planning fails when it is done from the LIMS database size. The transactional record is tiny; the raw instrument output is enormous, and it is the raw output that retention rules protect.
| Instrument class | Raw output | Volume per unit of work |
|---|---|---|
| LC-MS / GC-MS | Proprietary raw plus centroided peak lists | 1–8 GB per 24 h acquisition |
| High-throughput sequencer | Base-call files converted to FASTQ | 100–400 GB per run, several runs per week |
| Digital pathology slide scanner | Pyramidal whole-slide image | 1–4 GB per slide; 500 slides/day is 0.5–2 TB/day |
| Cryo-EM / TEM | Frame stacks and tilt series | 1–10 GB per tilt series, terabyte-scale per project |
| Spectroscopy / XRD | Small binary plus PDF reports | 10–500 MB per day |
| LIMS transactional database | Structured records and audit trail | Under 50 GB — but the most regulated data in the building |
| Instrument control PC | Operating system plus local acquisition cache | 500 GB–2 TB NVMe per bench |
The arithmetic surprises people. A single high-throughput sequencer at 300 GB per run and four runs a week produces roughly 62 TB of raw data a year before derived files are counted. Add one slide scanner at a terabyte a day and the lab is planning a petabyte-scale archive, not a server upgrade.
Retention is set by the regulation the lab operates under, and the periods differ by an order of magnitude. The correct design question is not "how much disk" but "which records, for how long, and what must the archive be able to prove".
| Regime | Scope | Retention floor | What the archive must prove |
|---|---|---|---|
| FDA 21 CFR Part 11 | Electronic records and signatures | Typically 5 years after last use; FDA may require longer | Audit trail, no silent edits, bound electronic signatures |
| EU GMP Annex 11 | Computerised systems in GMP | Product shelf life plus one year, minimum 5 years | Validated system, data integrity, change control |
| ISO/IEC 17025 | Testing and calibration laboratories | Commonly 6 years, set by the accreditor | Traceability and retrievability of the original observation |
| OECD GLP | Regulatory safety studies | 15–60 years depending on product type | A controlled archive under a named archivist |
| CLIA (42 CFR 493.1105) | US clinical laboratories | 2 years, longer for some record types | Test reports and quality records retrievable on demand |
| HIPAA / GDPR | Personal health and personal data | Jurisdiction-specific, sometimes shorter than the lab wants | Encryption, access control and a defensible erasure path |
Two consequences follow immediately. First, retention and privacy can pull in opposite directions: data that must be kept for twenty years under GLP may also contain personal identifiers that must be minimised. Pseudonymise at acquisition, not at archive. Second, "the derived file is enough" is not a defence — under OECD GLP principles the raw data must be retained, and derived results cannot substitute for it.
Data integrity guidance (ALCOA+: attributable, legible, contemporaneous, original, accurate, plus complete, consistent, enduring and available) maps directly onto a four-tier storage design.
| Tier | Technology | Purpose | Specification that matters |
|---|---|---|---|
| Instrument-local cache | 1–2 TB NVMe M.2 or U.2 | Acquire without dropping a frame when the network falters | Power-loss protection, high write endurance, checksums, wide temperature |
| Bench working store | 8–24 TB SSD in mirror or ZFS RAID | Live analysis, scratch space, derived files | End-to-end checksums, snapshots, scheduled scrub |
| Immutable lab vault | HDD RAID plus an object or WORM layer | Retained raw data with a hash chain | Write-once retention, object lock, per-file manifest hash |
| Offsite archive | Object storage or tape, geographically separate | Disaster recovery and long-horizon retention | Bucket-level immutability, restore testing |
| LIMS database | NVMe mirror with point-in-time recovery | Transactions, audit trail, e-signatures | Low capacity, highest integrity, restore drills |
Keep the vault write-once and keep the working store disposable. A lab that stores the only copy of its raw data on the analysis server has no archive; it has a single point of failure with a compliance label attached. Verify the chain, too: a periodic checksum scrub over the vault is the only mechanism that catches silent corruption before an inspection does, and a hash manifest stored alongside the data is what turns a folder of files into a defensible record.
Sizing follows from daily raw ingest rather than headcount. Take the sustained ingest, multiply by 250 working days, then apply the retention horizon: a lab taking in 250 GB a day is accumulating roughly 62 TB a year, so a ten-year obligation over the same bench is 620 TB before the derived analysis files that commonly double it. Budget the vault against that number and budget the bench cache against a single day of output plus the longest network outage the lab is prepared to tolerate — not against whatever capacity the instrument vendor's default configuration happens to ship with.
Planning a lab data platform or an instrument refresh?
Give us your instrument list, daily ingest and retention obligation — our engineers return a tiered storage BOM with the endurance, immutability and audit requirements budgeted per tier.
Contact: +86 137-1464-6179 | info@qscompute.com