01

Availability is a system property

Servers, networks, power, identity, applications, vendors, and people form one service chain. A component may be healthy while the service remains unavailable.

Start with essential services and work backward: which resources do they require, which dependencies are shared, and how long can the organization operate without each one?

  • Power and cooling
  • Network and internet access
  • Identity and authentication
  • Compute and storage
  • Backup and recovery
  • People and vendors
  • Documentation and communications
02

1–3: single points, capacity, and obsolescence

The first risk is a single point of failure: one link, device, credential, or specialist whose loss interrupts the service. The second is operating at the limit with no margin for peaks, failures, or maintenance. The third is retaining unsupported components with uncertain replacement or remediation.

Real redundancy requires independent paths and testing. Two devices connected to the same circuit, provider, or incorrect configuration can fail together.

03

4–5: untested backups and unmapped dependencies

Backup is not recovery. Copies require protection, appropriate retention and isolation, and restoration tests that prove timing and integrity. RPO defines acceptable data loss; RTO defines acceptable service downtime.

Hidden dependencies emerge when an application relies on DNS, directories, certificates, external APIs, or a key person and those relationships are undocumented. Service maps and runbooks reduce discovery time during a crisis.

04

6–7: fragile changes and improvised response

Changes without impact review, a defined window, validation, and rollback turn maintenance into incidents. Improvised response expands damage as outdated contacts, unclear authority, and inconsistent communications consume critical minutes.

A useful plan defines roles, triggers, alternate channels, recovery sequence, and closure criteria. It must be exercised, not merely stored.

05

Turn risk into a prioritized plan

Record each risk with the affected service, likelihood, impact, current control, owner, and target date. Prioritize the combination of high impact and weak recovery capability.

Use the NIST CSF 2.0 functions — Govern, Identify, Protect, Detect, Respond, and Recover — as common language between leadership and technical teams. The framework organizes outcomes; implementation must fit the organization.

Practical application

Operational readiness test

Any uncertain answer already indicates work to be done:

01Do critical services have an owner, RTO, and RPO?02Are single points of failure documented?03Is capacity margin and trend monitored?04Are backups restored in recurring tests?05Are emergency credentials protected and available?06Do changes include rollback procedures?07Are critical vendor contacts current?08Has the plan been exercised in the past 12 months?

Related technical references

External links to official sources. Always consult the current version and your organization’s specific context.

Informational content. It does not replace a specific technical, legal, or regulatory assessment.