For continuity and resilience managers
Recovery objectives are commitments, not achievements, until somebody has actually restored from a copy and timed it. This page states the objectives, says how the restore is exercised, describes what keeps running while it is degraded, and names the dependencies whose own status page will read healthy while your operation is not.
Continuity work has a peculiar shape: the artefact everybody asks for is a document, and the document tells you almost nothing. A binder with recovery objectives, a call tree and a dependency diagram is produced for an audit and satisfies it. Whether anybody has ever restored a copy and timed it is a completely separate question and it is asked far less often.
Rehearsal is where the truth lives and rehearsal is expensive, so it slips. The exercise gets scheduled, the quarter gets busy, the tabletop becomes a discussion instead of a drill, and the discussion concludes that the plan is sound. Everyone present knows the difference and nobody has the budget to say so.
Dependencies are where the real surprises are. The service that is down is rarely the one whose status page is red — an authentication provider, a certificate authority, a name resolution service, a message queue, a shared library’s package host. During an event, half the effort goes into establishing what is actually broken while several suppliers report themselves healthy, truthfully, because their own component is.
Degraded operation is the capability nobody specifies and everybody needs. A binary view of a service — working or not — produces plans in which the only response to a fault is to restore. What an operation actually needs is to keep moving with less: a queue that keeps accepting, a decision that still gets made on a computed rule, a record that gets written and reconciled later.
And there is the question your executives will ask afterwards, which is not what your recovery objective was. It is how long the operation was actually impaired, what the people doing the work did during it, and what was lost between the last good copy and the fault. Those three are the ones that matter and they are rarely the ones the plan is written to answer.
A recovery objective describes how quickly a service returns, and says nothing about whether the work stops in the meantime — so an operation whose only response to a fault is to restore has no answer for the hours before the restore completes.
The design intent that matters most here is that the operating decisions do not depend on the least reliable components. Routing, scheduling, ordering, thresholds and the rules that determine what happens next are computed rather than generated, so a model provider that is unreachable degrades the quality of an output rather than stopping the operation. That is the same property described on the provenance page and it is doing continuity work here.
Intake keeps accepting during a partial fault rather than refusing. Work that arrives while a downstream component is impaired is captured, held and reconciled when the component returns, so the failure mode of a partial outage is a backlog rather than a set of lost requests. A backlog is recoverable and a refused request is somebody who went elsewhere.
Backups are taken on a fixed cycle and held in the same region as the primary, which is a deliberate residency trade rather than an oversight, and it is stated on the residency page as well because it belongs in both decisions. The restore is exercised rather than assumed, and the number worth asking for is when a restore was last performed and how long it took — not what the objective says.
The honest limit is here rather than at the bottom: a full regional loss of the hosting provider is a longer event than a component fault, and the recovery position in that case is a restore in another region rather than a running standby. For an organisation whose obligation requires a warm standby in a second region, this arrangement does not meet it today, and that is a decline rather than a negotiation.
The dependency list is given specifically rather than in categories, because the whole difficulty during an event is establishing what is broken while everybody reports healthy. Knowing in advance which providers sit underneath, and what each one being unavailable actually does to the operation, is the artefact that makes an event shorter.
When a restore was last actually performed — measured by a date and a duration rather than an objective from a plan document.
What keeps working during a partial fault — measured by disabling a dependency in a trial and observing which operations continue.
Whether intake refuses or holds — measured by submitting work while a downstream component is unavailable and confirming it is captured.
Whether held work reconciles — measured by restoring the component and confirming the held work completes rather than being discarded.
Which dependencies would report themselves healthy while you are impaired — measured by the written dependency list, with the effect of each one being unavailable.
What was lost between the last copy and the fault — measured by the backup interval, stated as a number, and confirmed by a restore exercise.
no warm standby in a second region is operated today; a regional loss is a restore into another region and a materially longer event, and an obligation requiring a running standby is a decline. No independent audit of continuity arrangements has been performed, and no continuity certification is held or claimed. An independent SOC 2 Type II attestation is in progress and no report exists yet. Recovery objectives are commitments rather than a record of achieved performance.
Where connections into your estate run under credentials you hold, an event on our side does not take your provider relationships with it — your messaging, your payment processing and your model provider are your accounts and your contracts, and they continue to exist independently of what is happening to us.
Your log stream keeps its copy of everything that already happened, which matters more during an event than at any other time. An organisation whose only record of the last month lives inside an impaired supplier is in a considerably worse position than one holding its own copy, and the difference is invisible until precisely this moment.
And the export described on the portability page is the ultimate continuity control. A recent export in your possession means the worst case is a slow recovery rather than a lost operation, and taking one on a schedule is the cheapest resilience measure available to any customer of any supplier.
The most useful sentence on any continuity page is the one naming what has never been rehearsed, because every plan has one and most pages omit it. Here it is the full regional case: a loss of the hosting region is a restore elsewhere rather than a failover, it is a materially longer event than a component fault, and it has not been exercised end to end at production scale.
What is exercised is the restore from a backup copy, and the number worth asking for is the date of the last one and the duration it took rather than the objective in a document. Ask every supplier for that pair. A supplier who can only produce an objective has told you the plan is the artefact.
Degraded operation is designed rather than attested, and it is the property most worth testing yourself. Disable a dependency in a trial and watch what continues. That single exercise tells a resilience manager more than any document, and it is the test that distinguishes an operation that keeps moving from one that only has a restore.
No continuity certification or independent audit of these arrangements exists. A SOC 2 Type II attestation is in progress and no report exists yet; availability would fall within the scope of the one under way rather than being a separate exercise, and it is an attestation with a defined period rather than a certification.
When was a restore last performed, and how long did it take. Ask every supplier this pair, in writing, and compare the answer to their stated objective. A supplier who answers with an objective has answered a different question, and a supplier whose last restore was long ago has a plan rather than a capability.
What continues while it is impaired. This is the question that determines what your own operation looks like during an event, and it is almost never in a questionnaire. The answer should be specific — which operations continue, which degrade, and which stop — rather than a general assurance about resilience.
Then read the dependency list for the effect of each one being unavailable rather than for the names. During an event, the difficulty is establishing what is broken while several suppliers truthfully report healthy, and a list of names does not shorten that.
And take your own export on a schedule regardless of what any supplier commits to. It is the cheapest resilience measure available to you, it works no matter which supplier is impaired, and it converts a worst case from a lost operation into a slow one.
One trial in which a dependency is deliberately disabled and your team observes which operations continue, which degrade and which stop — before any continuity documentation is exchanged.
A discussion produces agreement and an exercise produces information. Disabling one dependency in a trial takes an afternoon and answers the question your tabletop is actually about, which is what your own operation looks like during somebody else’s event.
Submit work while the dependency is down and confirm it is held rather than refused, then restore the dependency and confirm the held work completes. That pair is the whole continuity argument on this page, and it is checkable rather than assertable.
Then ask for the last restore date and duration, and read the dependency list for effects. Those two artefacts, plus the exercise you just ran, are a better continuity file than any binder either party could produce.
They are commitments and they are the less useful half of the answer, so ask the other half in the same breath: when was a restore last actually performed and how long did it take. A plan document produces objectives for an audit and tells you nothing about capability. Ask every supplier for the date and the duration — the ones who can only give you an objective have told you which artefact they maintain, and the gap between the two is where continuity programmes fail.
No. There is no warm standby in a second region today, so a regional loss of the hosting provider is a restore elsewhere and a materially longer event than a component fault. If your obligation requires a running standby, this arrangement does not meet it and that is a decline rather than something to negotiate. It is stated here rather than discovered in your tabletop, which is where a resilience manager would otherwise find it and rightly be annoyed.
That is the right question and it is almost never in a questionnaire. Operating decisions are computed rather than generated, so a model provider being unreachable reduces output quality without stopping the work. Intake keeps accepting while a downstream component is impaired, holding what arrives and reconciling it afterwards — so a partial outage produces a backlog rather than lost requests. Test both in a trial by disabling a dependency; an afternoon of that is worth more than any document either of us could write.
They usually are, which is why the dependency list is given with the effect of each one being unavailable rather than as a list of names. During an event the expensive part is establishing what is actually broken while several suppliers truthfully report themselves healthy, because their own component is. Knowing in advance which providers sit underneath, and what the operation does when each is unreachable, is the artefact that shortens an event — and it is one you should demand from every supplier in the chain, not only from us.
No. No continuity certification is held and no independent audit of these arrangements exists. A SOC 2 Type II attestation is in progress, no report exists yet, and availability would fall within its scope rather than being a separate exercise — but it is an attestation with a defined period rather than a certification and nobody can be handed one today. Record the degraded-operation properties as designed rather than attested in your file, and close the gap yourself with the dependency-disabling exercise, which tests the property directly.