Edge Site Backup & Disaster Recovery 2026:
3-2-1-1-0, RPO/RTO & Storage for Unattended Sites

Published: September 18, 2026 | Category: Buying Guide | QSCompute

Four hundred sites, no IT staff, a 20 Mbps uplink and a power cut every storm season. That is where most industrial edge data now lives, and it breaks almost every assumption a data-centre backup design rests on: nobody is there to swap media, the network cannot carry the raw data, and a site may be dark for hours while the line keeps producing records that will later have to be defended.

This guide covers how to classify edge data by recovery point and recovery time, how to apply the 3-2-1-1-0 rule when the uplink is the binding constraint, and what to specify in the storage and compute hardware that implements it. The recurring theme is that backup is cheap and restore is expensive, so design and budget for the restore.

Classify the Data Before Choosing Hardware

The first mistake at an edge site is treating all data as one pool. An inspection archive and a machine recipe have nothing in common: one is large, replaceable and rarely read after thirty days, the other is tiny, unique, and the reason the line cannot start after a failure. Recovery point objective and recovery time objective — how much data you can lose, and how long you can be down — are the two numbers that separate them, and they drive the storage tier each class deserves.

Data classChange rateRPORTORetentionWhere it should live
Controller and IPC configuration, recipes, calibrationOn changeZero — version every writeMinutesFull historyLocal SSD plus offsite, versioned as text
Trained models and model artefactsPer releaseZero at releaseMinutesAll released versionsLocal plus immutable offsite; never store only on the site
Operating-system and container imagesPer patch cycleOne version backUnder an hourTwo to three versionsLocal image store, rebuildable from the registry
Regulated records: batch, quality, audit trailContinuousUnder an hourHours (offline acceptable)Regime-definedAppend-only store plus offsite copy with object lock
Telemetry and historian dataContinuous, high rateMinutes to hoursHoursLocal window, then aggregateLocal time-series store; ship roll-ups, not raw samples
Video and vision evidenceVery highHoursDaysDays locally, longer on eventsRolling local buffer; ship events, incidents and clips only
Debug logs and tracesContinuousBest effortNot applicableDays, cappedSeparate partition with a hard size cap

The order matters. Configuration, recipes and model artefacts are usually the smallest classes on site and the ones that stop production when lost, yet they are the least protected because they never look like the big data problem. The video archive, which is genuinely large, is the class whose recovery matters least — which is why it belongs on a rolling local buffer.

The 3-2-1-1-0 Rule at an Edge Site

The classical rule asks for three copies of the data, on two different media, with one copy offsite. The modern extension adds a fourth copy that is offline or immutable, and a zero standing for verified — the restore is tested and the check reports success, rather than being assumed. At an edge site each of those five elements has a hardware implication.

Three copies means the production device, a local second copy and an offsite copy. Two media means the local second copy should not share the same failure domain — a mirrored pair of identical drives in one enclosure survives a drive failure, not a power event or a firmware bug. The immutable copy answers ransomware and compromised credentials: write-once retention, or a store that is not mounted and not reachable from the site. The verified zero is a scheduled restore, not a scheduled backup.

The uplink forces the design. Eight cameras at 20 GB/day each produce 160 GB/day of raw video, which no 20 Mbps link carries. The practical answer is to ship deltas and derived data: configuration and model artefacts are kilobytes to megabytes and go on every change, telemetry ships as compressed roll-ups, and video ships as event clips with the rolling buffer retained locally. Block-level incremental tools make the difference concrete — a full image of a 200 GB node is 200 GB once, then the daily delta is typically single-digit gigabytes, which fits an overnight window on a constrained link.

TierScope and methodTypical size per siteUplink needRestore path
Local liveProduction SSD in the node; RAID or mirror for the record storeTens to hundreds of GBNoneIn place, minutes
Local snapshotFilesystem or volume snapshots on a second device, hourly to dailyDelta onlyNoneRollback in minutes; recovers from operator error
Site storeSmall NAS or industrial edge server holding image, config and clip copies0.5–4 TBNoneBare-metal restore over the site LAN
Offsite incrementalBlock-level incremental push to central storage or object storeSingle-digit GB/day deltaFits an overnight window on 10–20 MbpsPull over WAN, or ship an encrypted drive for full restore
Immutable / air-gappedObject lock with a retention period, or an offline removable copyConfig, records, model artefactsLowRansomware and credential-compromise recovery

Two operational details decide whether this runs. First, store-and-forward: telemetry and record traffic must queue locally and survive a multi-day outage without loss or duplication, so the local queue has to be persistent rather than an in-memory buffer. Second, cap the growth — an uncapped log partition will eventually fill the device that also holds the operating system.

Restore Is the Only Metric That Counts

A backup that has never been restored is a hypothesis. The failure modes at an edge site are specific. Secure boot with keys sealed to a TPM is the first: if the platform seals its measurements to the exact firmware and configuration state, a restored image may refuse to boot on a replacement node, so recovery keys must be escrowed where the site cannot lose them. The second is consistency: copying a live database or time-series directory produces a file that opens and is wrong, so either stop the service and snapshot, or use a crash-consistent filesystem or volume snapshot.

The third is arithmetic. A restore that has to pull a full multi-hundred-gigabyte image over 20 Mbps takes days, not hours, so the plan needs a cold spare at the site or a regional depot with the image staged, plus an encrypted physical copy for the worst case. Rehearsing that restore on a bench is the only way to find the missing step before a storm night forces the discovery.

RequirementWhat to specify
Local redundancyTwo independent devices or a mirrored volume for the record store; RAID is not a backup
SnapshotsFilesystem or volume snapshots with a bounded retention policy and a scheduled prune
Offsite transportBlock-level incremental or content-addressed deduplication; resumable transfers; bandwidth limiting outside production hours
ImmutabilityObject lock or write-once retention on at least one offsite copy; credentials for that copy deliberately unavailable to the site
EncryptionEncryption in transit and at rest, with keys held centrally and recovery keys escrowed outside the site
ConsistencyService stop-and-snapshot or a crash-consistent snapshot; application-consistent hooks where a database is involved
Bootability of a restoreDocumented path for TPM-sealed secure boot: key escrow, re-provisioning, or a golden image that reseals on first boot
Uplink realismBulk transfer budget sized to the site's real sustained uplink, not its nominal line rate
Storage enduranceIndustrial SSD with power-loss protection and DWPD sized for the backup write amplification on the receiving device
TelemetryBackup success and restore-test results reported into the same fleet monitoring used for the machines

Five rules for a defensible edge backup design:

  1. Classify first, buy second. RPO and RTO per data class tell you which tier each dataset needs; buying one storage tier for everything guarantees you overpay for video and under-protect recipes.
  2. Ship deltas and derived data. Raw video over a constrained uplink is a bandwidth decision nobody can justify; events, clips, roll-ups and block-level increments are not.
  3. Keep one copy the site cannot reach. Immutability is what survives ransomware and stolen credentials, and it is cheap at the sizes the critical classes occupy.
  4. Rehearse the restore, including the boot. A TPM-sealed platform that will not accept a restored image is the most common nasty surprise, and it is found on a bench or not at all.
  5. Cap every partition. Logs and backups need hard size limits so that the protection mechanism cannot itself fill the device.

QSCompute supplies the hardware layer of edge data protection — fanless industrial PCs and edge servers for the site store, industrial SSD and NVMe with power-loss protection for record and backup volumes, U.2 and E.1.S drives for RAID tiers, plus industrial DRAM, wide-input DC supplies and rugged NAS-ready chassis. DDP shipping to 85+ countries.

Designing backup and recovery for unattended edge sites?

Send us your site count, per-site data classes, uplink capacity and required RPO/RTO — our engineers return a tiered storage and compute BOM with the offsite bandwidth budget and restore path worked out.

Contact: +86 137-1464-6179 | info@qscompute.com