Resilience, backup, and recovery
Test restoration of the required system state and plan operation during dependency outages. Include reconciliation of changes made while services were unavailable.
Sources and scopeSource record 25 August 2026
Technical source record: 25 August 2026. Check the linked documentation for current product requirements.
Research and static guidance only; product and deployment specific behaviour requires controlled environment validation and authoritative product evidence.
On this page
Overview#
Availability isn't a single uptime percentage. A physical security system must define what each edge device, controller, server, integration and operator can safely do when dependencies fail, and how authoritative state is reconciled after recovery.
Dependency failure matrix#
Evaluate loss of power, network path, DNS, DHCP, time, directory/IdP, CA/revocation, broker, database, storage, management server, controller, cloud, cellular path, primary site and operator workstation. For each dependency record:
- locally retained capability and duration;
- default physical behaviour and safety authority;
- queue/storage capacity and overflow behaviour;
- operator indication and monitoring evidence;
- retry/backoff and failover sequence;
- split brain or duplicate command risk;
- recovery ordering and state reconciliation;
- conditions requiring qualified/manual intervention.
Backup scope#
Back up more than application databases: device/controller configuration, system topology, identities and roles, schedules and rules, certificates and trust anchors, encryption/signing key custody information, licence/entitlement data, integration mappings, broker configuration, evidence indexes, firmware/software installers, documentation, and recovery credentials.
Keep keys and secrets under appropriate separate protection; a configuration backup that can't be decrypted or a backup that exposes every private key isn't a workable recovery design.
Backup properties#
- defined source and consistency point;
- authenticated and encrypted transfer;
- immutable or separately administered copy;
- retention and deletion aligned with sensitive data policy;
- inventory of dependencies needed to restore;
- integrity verification and anomaly monitoring;
- protected recovery credentials and break glass access;
- controlled restoration exercise with exact versions, accountable owner, and observed result.
Recovery sequence#
- Establish safety and preserve relevant evidence.
- Contain compromised trust, accounts, network paths and update sources.
- Choose a known good hardware, firmware, software and configuration baseline.
- Restore identity, trust, time and core infrastructure in a documented order.
- Restore management and integrations without immediately reconnecting untrusted devices.
- Re enrol or rekey endpoints where trust may be compromised.
- Reconcile queued events, commands, credentials, schedules, recordings and controller state.
- Complete and approve functional, failure, audit, and safety validation before returning to normal service.
Sources#
- NIST 800 82, NIST SP 800-82 Rev. 3, contingency and OT resilience guidance, accessed 25 August 2026.
- NIST 1339, NIST SP 1339: Operational Technology Backup Quick Start Guide, final June 2026, accessed 25 August 2026.