Operational resilience is the ability to keep delivering the services that matter when the resources those services depend on are impaired. It is an outcome standard, not a document type. A business continuity plan describes how you respond. Operational resilience asks a harder question first: which services must keep running, how long they can be degraded before the harm is intolerable, and whether you have actually proved you can stay inside that limit.
Most programs fail this the same way. They inventory departments, write recovery procedures for systems, and call the result resilience. The customer experiences a service — a payment that clears, a claim that is paid, a record that opens — not a department. This guide is the sequence for building that service view, setting tolerances you can test, and governing the resources, including third parties, underneath.
What operational resilience is, and what it is not
The service outcome
What it means: Operational resilience starts from the external outcome. An important business service is a service whose disruption would harm customers, the firm, or — in regulated markets — financial stability or market integrity. Internal activities that exist only to support that outcome are resources, not services. Payroll may be critical to the firm and still not be an important business service in this sense. The distinction matters because the tolerance, the mapping, and the test are all built on the service, not on the org chart.
What good looks like: A short, named list of important business services, each written as an outcome a customer or market participant would recognize. Each service has one accountable owner, a plain-language description of who is harmed if it stops, and a boundary that says what is inside the service and what is a dependency. The list is stable enough to govern and short enough that leadership can recite it.
Common failure mode: The “important services” list is the department directory with the word “service” added. Forty entries, no owner, no harm statement. Everything is critical, so nothing is prioritized, and the test program never reaches the services that would actually hurt customers.
Not a second business continuity plan
What it means: Business continuity management — the plan, the business impact analysis, the recovery strategies, the exercises — is how you build the capability. Operational resilience is the standard you hold that capability to: stay inside an impact tolerance for the important services, including when the disruption is severe. ISO 22301 remains the management-system language many firms already use for continuity. Operational resilience does not replace it. It changes the question the management system has to answer, from “did we recover the process” to “did the important service stay within tolerance.”
Common failure mode: A parallel “operational resilience” workstream that rewrites the same plans in new templates, owned by a different team, with no authority to change the recovery design when a test shows the tolerance cannot be met.
Identify the services before you map the machinery
What it means: Identification is a judgment about harm, not a scoring model that produces a number nobody will defend. For each candidate service, write who is harmed, how quickly, and whether a substitute exists that the customer can actually use. Include services delivered through a third party if the customer still experiences them as yours. Exclude internal conveniences that can stop for days without that harm.
What good looks like: A workshop with the people who run the service and the people who face the customer, facilitated by someone who will reject vague names. Each service gets a one-page charter: outcome, customers, hours it must be available, owner, and the harm that defines “intolerable.” Disagreements are recorded. The charter is approved by someone who can later fund a gap.
Common failure mode: A consultant questionnaire scored in a spreadsheet. The output is a heat map. No one in the business recognizes the service names, and the owner box says “TBD” six months later.
Map the resources the service actually uses
What it means: A dependency map for an important business service is the chain of people, premises, technology, data, and third parties required to deliver the outcome — not the architecture diagram of one application. Follow the service through a normal day and through the ugly cases: end of month, a peak trading day, a night shift, a manual workaround. Note where the chain leaves the firm.
What good looks like: For each service, a map that names the resources at a level you can test. “Payments platform” is not a resource; the specific processing path, the identity provider it calls, the scheme it settles through, and the operations team that clears exceptions are resources. Each resource has a stand-in or it is marked as a single point. The map is updated when the service changes, not on an annual anniversary nobody attends.
Common failure mode: A technology inventory labeled as a dependency map. People, premises, and vendors are missing. The identity service, the DNS path, and the licensing server that the “resilient” application quietly requires are discovered during the incident, which is the most expensive way to draw the map.
Set impact tolerances you can fail
Tolerance is a harm threshold, not a recovery target
What it means: An impact tolerance is the maximum level of disruption to an important business service that the firm can tolerate before the harm — to customers, to the firm, or to the wider system — becomes intolerable. A recovery time objective is a design target for how fast you intend to restore a process or a system. They should be consistent. They are not the same sentence. You can meet every RTO in the business impact analysis and still breach an impact tolerance if the service is degraded in a way the RTO never measured: wrong data, partial function, a channel down while the core is “up.”
What good looks like: Each important business service has a tolerance stated in time and in harm: how long a full outage is tolerable, how long a defined degradation is tolerable, and what “intolerable” means in customer or market terms. The business impact analysis supplies the recovery priorities underneath — which processes and systems must be restored in which order to stay inside that tolerance. The tolerance is approved by someone who will accept the test result when the service misses it.
Common failure mode: The impact tolerance is copied from the shortest RTO in the business impact analysis, or it is set at a number that has never been missed because the scenarios are too kind. A tolerance that cannot be breached is not a tolerance. It is a slogan.
Use the business impact analysis for the priorities underneath
What it means: The business impact analysis still does the work of ranking processes, estimating how impact grows with time, and setting recovery time and recovery point objectives for the resources. Operational resilience does not skip that work. It stops the firm from treating those objectives as the finish line. Recovery priorities exist so the important service stays inside tolerance. If the priorities and the tolerance contradict each other, the tolerance wins and the recovery design changes.
What good looks like: A traceable line from service, to tolerance, to the processes that deliver it, to the RTO and RPO of each resource, to the strategy that makes those numbers real. A named person signs the business impact analysis. A model, a dashboard, or a copied prior year does not.
Common failure mode: Recovery objectives negotiated with IT to match what the current backup already does. The tolerance is then written to fit the backup. The customer harm was never the input.
Test severe but plausible scenarios against the tolerance
What it means: A resilience test asks whether the important business service stays inside its impact tolerance when a severe but plausible disruption hits the resources it depends on. Severe means the disruption is bad enough to threaten the tolerance. Plausible means it could happen to this firm, given its dependencies, not that it is the most likely Tuesday. The point of the test is to find the breach while you can still change the design.
What good looks like: A small number of scenarios per service, each aimed at a real dependency: loss of a third party that cannot be swapped in the tolerance window, loss of a site and the people in it, corruption or loss of the data the service cannot reconstruct, a cyber event that removes both the primary path and the usual workaround. The scenario is run far enough to see whether the service is actually delivered, not merely whether the crisis team convened. Findings change the map, the strategy, or the tolerance. An after-action report with no change to any of the three is incomplete.
Common failure mode: An annual tabletop that confirms the call tree works, scored as a passed resilience test. Discussion is a legitimate exercise type. It is not, by itself, evidence that the service stayed inside tolerance. A test that cannot fail has not tested anything.
Third parties sit inside the tolerance
What it means: If an important business service depends on a supplier, a cloud platform, a payments scheme, or a subcontractor of your supplier, that party is part of the resilience of the service. Their contract, their own recovery, and your ability to exit or substitute them have to be judged against the same impact tolerance. Mapping only your tier-one vendor, and stopping there, leaves the chain that will actually break untested.
What good looks like: For each important service, the third parties on the critical path are named, including the ones your vendor uses where you can see them. You know which of them you can replace inside the tolerance and which you cannot. Where you cannot replace them, the vulnerability is explicit: concentration, exit plan, and the scenario that assumes they are gone. Supply-chain work — second sources, buffers, and rehearsed breaks — is funded against that list, not against a generic vendor-risk score.
Common failure mode: A vendor questionnaire filed as resilience evidence. The questionnaire says the vendor “has a business continuity plan.” Nobody has asked whether that plan restores the specific service inside your tolerance, and nobody has mapped the vendor’s own single points.
Governance: a named owner, a change trigger, a board view
What it means: Operational resilience fails in the gap between the team that writes the map and the executive who can spend money when the map shows a breach. Governance is the set of decisions that close that gap: who owns each important business service, who accepts a tolerance, who receives a failed test, and what change in the business forces the map to be redrawn.
What good looks like: Each service has a named owner who is still in the role, not a job title that was reorganized away. The board or its risk committee sees the service list, the tolerances, the last test result, and the open vulnerabilities — not a maturity score. Change triggers are written down: a new product, a new supplier on the critical path, a move of the people who run the service, a failed test. Any one of those reopens the charter.
Common failure mode: Ownership assigned to the resilience team. The resilience team can document the gap. It cannot change the architecture, the contract, or the staffing model. The vulnerability is reported every quarter and funded never.
A working sequence
- Name the important business services as customer outcomes, with a harm statement and a single owner.
- Map the people, premises, technology, data, and third parties each service needs, including the shared services everyone forgets.
- Set an impact tolerance in time and in harm. Use the business impact analysis to set the recovery priorities that keep the service inside it.
- Run a small number of severe but plausible scenarios aimed at the dependencies that would breach the tolerance. Require the test to be able to fail.
- Change the design, the contract, the staffing, or the tolerance. Record which one moved.
- Reopen the work when the service changes, not when the calendar says the annual review is due.
FAQ
Is operational resilience the same as business continuity?
No. Business continuity is the management discipline: analysis, strategies, plans, and exercises. Operational resilience is the outcome those activities are supposed to produce for the services that matter — staying inside an impact tolerance when disruption is severe. Firms that already run a continuity program should connect it to services and tolerances, not start a second program beside it.
How is an impact tolerance different from an RTO?
An impact tolerance is the point at which disruption of an important business service causes intolerable harm. A recovery time objective is how fast you intend to restore a given process or system. The RTO is a design input. The tolerance is the harm limit the design has to respect. Meeting the RTO does not prove the tolerance was met if the service was degraded in a way the RTO did not measure.
How many important business services should a firm have?
Few enough that each one can be mapped, owned, and tested. If the list is long enough that no scenario ever reaches most of it, the list is a catalog, not a resilience scope. Split or drop entries that are internal activities rather than customer or market outcomes.
Does a tabletop exercise satisfy operational resilience testing?
A tabletop is discussion. It is useful for roles, decisions, and finding the first gaps. It does not, by itself, show that an important business service stayed inside its impact tolerance. Use it to design the harder test. Do not file it as the harder test.
Where do third parties belong?
On the dependency map of the service, judged against the same tolerance as internal resources. A vendor business-continuity attestation is not a substitute for knowing whether you can deliver the service if that vendor stops.
How this connects
Identification, mapping, and impact tolerances are worked in detail in the Important Business Services guide. Scenario design against severe but plausible disruption is in the Operational Resilience Testing guide. The supplier and concentration side is in the Supply Chain Resilience guide. Recovery priorities underneath the tolerance are built with the Business Impact Analysis program guide and the method for building a BIA and recovery priorities — a method, not a software tool. Where rules are stacking on the same firms, the regulatory convergence briefing separates DORA, CISA reporting, and ISO 22301 instead of treating them as one standard.