Continuity When the Model Goes Down
The short answer
Every AI-dependent process has a default continuity plan already, whether anyone wrote it down or not
If nobody classified the process and designed a degraded mode, the default is silence, discovered by whoever is on shift when the outage happens. This topic replaces that default with a deliberate one.
What you will be able to do
- Classify every AI-dependent process in your organization into one of three continuity classes: pausable, human-fallback, or cannot-stop, using a repeatable test rather than intuition.
- Design a degraded mode for each class, stating in plain language what the process does the moment the model becomes unavailable, and who or what performs that function instead.
- Set a recovery objective for each process that the business, not the technical team alone, has explicitly signed, naming the maximum tolerable outage window and the maximum acceptable staleness of the last AI output the process relied on.
- Analyze a real, documented cloud-provider outage (the Microsoft Azure Front Door incident of 30 July 2024) for the specific mechanism that turned a contained attack into an eight-hour, cross-service disruption, and connect that mechanism to a defense your own continuity plan must not repeat.
- Schedule a failover test on a fixed cadence, and distinguish a genuine test (the fallback is actually exercised) from a paper test (the fallback is described but never exercised).
- Build a continuity table for five AI-dependent processes, each row carrying a classification, a degraded mode, a recovery objective, and a last-test date, defensible to a chief risk officer or a hostile board member line by line.
- Evaluate a vendor's incident-notification and model-change clauses (see Topic 8.7) against your own continuity table, and state exactly where the contract's promise ends and your own plan has to begin.
- Distinguish a single point of failure that lives inside your own organization from one that lives inside a shared upstream dependency (a cloud region, a content delivery network, a single model provider), and explain why the second kind cannot be engineered away, only planned around.
- Prioritize a limited continuity budget across five processes, defending which two or three earn a hot, tested failover and which are defensible with a documented pause and an honest customer message instead.
- Detect partial degradation, not only total outage, by naming the specific monitored signal and threshold that should trigger a degraded mode before the failure becomes undeniable to everyone downstream.
- Compare a rollback (undoing the organization's own recent release) against a degraded mode (absorbing a vendor's outage) (see Topic 3.4), and explain why an organization needs both rather than treating either as a substitute for the other.
The lesson
At 11.45 Coordinated Universal Time on July 30, 2024, a distributed denial-of-service attack hit Microsoft's Azure front door, the traffic layer sitting in front of a huge share of the internet's hosted applications. Ordinarily, this is a contained incident. The system built to absorb the punch instead swung with it.
An error in the implementation of the company's own defense mechanism amplified the impact of the attack rather than mitigating it. The operational result was an eight-hour blackout. The Azure portal would not reliably load, and a long list of Microsoft 365 services and downstream AI-dependent processes went completely dark.
During that window, organizations that believed they had built redundant, multi-vendor AI ecosystems watched all of their tools fail simultaneously. Infrastructure failure supersedes software layer redundancy. Catastrophic outages are an architectural inevitability, not a rare anomaly.
Every AI-dependent process in your organization already has a continuity plan for the moment the model stops answering. If nobody explicitly engineered that plan in advance, the default response is simply silence. Silence forces a scramble.
Business continuity decisions are made by accident, in the middle of a crisis, by whoever happens to answer the phone first. To replace that dangerous default, we build this, the continuity table. It is a deliberate, engineered framework for a targeted response.
Surviving a critical outage requires answering exactly what happens to each process before the crisis ever occurs. Column one is classification. We sort by the harm of a gap axiom.
The time before silence causes irreversible harm. Processes drop into three tiers. Pausable.
Waiting hours causes no damage. Human fallback. Manual teams absorb a wait of minutes to hours.
Cannot stop. Tolerance is seconds. Waiting causes immediate harm.
A quarterly strategy planning assistant is highly important to executives, but it is strictly possible. A low-profile, real-time fraud filter sitting in front of a payment flow is cannot stop. Confusing business importance with operational urgency guarantees you will misallocate resources exactly when you need them most.
Column two establishes the degraded mode. Vague intentions are crossed out, replaced by executable actions. For possible processes, the protocol uses a specific queue and message system.
For human fallback, it identifies exact personnel, dictates manual procedures, and sets a hard volume threshold. Outages generate their own surge in user demand. If you size a human fallback strictly for normal volume, the team will collapse under the predictable surge that accompanies the failure.
For a cannot stop process, humans cannot keep up. The protocol requires a redundant technical path, like a rules-based threshold check running on separate infrastructure. A classification label without an executable specific degraded mode is merely an opinion.
It is not an operational plan. Column three adds recovery objectives. The table expands to include two specific metrics, capped by a non-technical business owner's signature.
The first is maximum tolerable outage, defining how long the business tolerates degraded mode. If IT invents an RTO in isolation, business leaders will overrule it during a crisis. The second metric is maximum acceptable staleness, or RPO.
It defines how old a cached AI output is allowed to be before you must stop trusting it. Without an expiry limit, a stale cache becomes a severe liability. An aging output quietly transitions from a helpful placeholder into a confidently wrong decision.
You negotiate this recovery math during peacetime because establishing rational consensus is impossible under the of a live event. Column four tracks test cadence. This split-screen timeline illustrates the schedule.
Cannot-stop processes on top require frequent, tightly-packed tests, while pausable processes below need far fewer. A genuine test actually triggers the fallback under simulated load, rather than merely describing it. Remember the massive Azure front-door incident? Their defense failed because it behaved differently under actual hostile load than the engineers assumed it would behave on paper.
Exercising these fallbacks requires drawing from a finite continuity budget of engineering time. Testing every process uniformly across the organization wastes that budget on low-stakes tools and creates a false sense of security for the critical ones. You direct the continuity budget toward consequence.
You test cannot-stop processes heavily, and you test possible processes lightly. An untested fallback plan is simply a single point of failure hiding behind a compliance document. The final column forces the infrastructure check.
Looking at this network stack, five different vendor logos at the top funnel into a single cloud region. This is vendor concentration risk. Distinct tools provide zero redundancy if they route through the same foundation.
When it drops, they all drop. An enterprise cannot always engineer away this shared dependency due to budget constraints, but it is obligated to identify and price that risk openly. True architectural resilience requires mapping your dependencies down to the physical foundation, rather than stopping at the vendor logo layer.
Executives often assume that signing a strong vendor service level agreement guarantees operational resilience. An SLA only dictates what the vendor is required to do, namely notify you of an incident. It leaves your organization's own internal response completely undefined.
This stuttering mechanism represents a much more common reality, partial degradation. Real outages rarely present as a clean loss of signal. They present as latency spikes and intermittent errors.
You must establish defined degradation thresholds to trigger your fallback protocols. You trigger your continuity plans the moment those specific metrics cross the threshold, rather than waiting for an undeniable total failure to force your hand. Your immediate action item is to build a complete five-row continuity table covering your organization's core AI processes.
A regional insurer mapping a claims triage assistant as human fallback, while categorizing a fraud flagging model as cannot stop, drives vastly different engineering budgets for those two rows. Once built, the continuity table is a living artifact. It goes stale the exact moment an internal process or an upstream vendor changes.
You must mandate a reopening and revision of this table at every contract renewal, or whenever a material infrastructure shift occurs. Organizations survive major global outages by deciding in a calm room, weeks in advance, exactly what happens when the models go dark.
The ideas, one by one
Classify on the harm of a gap, not on the process's importance
A strategically important process can still be pausable if nothing time-critical happens while it waits; a low-profile process can still be cannot-stop if a short gap produces an irreversible harm. The two judgments are not the same question.
A degraded mode is a specific procedure with a named owner, never a vague intention
"The team will handle it" is not a plan. A colleague unfamiliar with the process should be able to execute the degraded mode from the written description alone.
Recovery objectives belong to the business, not the technical team alone
A number invented without the business owner's sign-off will not survive the first real incident; someone with more authority will overrule it on the spot, exactly when the plan needed to hold.
A described fallback and a tested fallback are different things
The Azure Front Door outage happened because a defense mechanism behaved differently under real hostile load than it was assumed to; only a genuine, triggered test closes that same gap in a continuity plan.
Different vendor names can still share one underlying point of failure
Check the infrastructure layer beneath each AI-dependent process, not only the vendor-facing product name, or a table that looks diversified can still fail as a single point of failure.
Test cadence should match continuity class, not a single organization-wide calendar
A cannot-stop process deserves frequent, disciplined testing; a pausable process does not need the same intensity, and treating both alike wastes the effort that the highest-stakes row actually needs.
A stale cached answer can be more dangerous than an honest gap
Set a maximum acceptable staleness for every process that relies on a cached or last-known-good output, and fail safe once it is exceeded, rather than serving an aging answer with no expiry.
The vendor's contract and your continuity table protect against different failures
A strong incident-notification clause tells you sooner that something is wrong (see Topic 8.7); it does not, by itself, tell your organization what to do about it. Both are required, and neither substitutes for the other.
A continuity table goes stale the moment the processes it covers change
Reopen it at contract renewal, at the first material change to any of its five processes, or on a fixed calendar, whichever comes first, the same discipline that keeps a model-change-notice clause meaningful over time.
Spend a limited continuity budget on consequence, not on visibility
The row where a gap cannot be undone earns the fullest, best-tested fallback, even if it is the least visible to a customer; the row a customer would notice first is not automatically the row that deserves the most protection.
A rollback and a degraded mode solve different problems
A rollback undoes your own organization's recent change (see Topic 3.4); a degraded mode absorbs a dependency you do not control going dark. An organization needs both, and neither is a substitute for the other.
You read it. Now prove it.
Explain this lesson in your own words, the way you would to a colleague, without looking back at it. It is graded against the lesson itself, by the same grader our learners face. One free try a day, no account needed.
The conversation
The same lesson, talked through at length by two hosts: the full transcript of the audio deep dive.
Listen to it as episode 67 of the podcast.
Read the full conversation
Imagine it is 11.45 a.m. on a Tuesday, you're sitting at your desk, maybe you're reviewing some quarterly projections, and suddenly your company's AI-powered customer service agent just stops. Just completely offline, absolute silence. Exactly, and this is a tool that routinely handles, you know, something like 10,000 queries an hour.
So because there is no backup plan in place, your human staff is instantly overwhelmed. Oh yeah, they are hit by a flood of confused customers. Millions of dollars in transactions are stalled, compliance checks are just bypassed, and frankly nobody on the floor knows exactly who is supposed to make the call on what to do next.
And I mean this isn't some dystopian hypothetical, this happens every single day in modern business. It really does, and the fallout from that specific moment is almost entirely dictated by decisions made months before the outage actually happened. Right, so welcome to the deep dive.
Our mission today is highly specific, and it is squarely aimed at you, the professional listening to this right now. We are going to construct a continuity table. Yes, a very precise mechanical framework.
Exactly, we're going to build the exact framework that dictates what your organization does the very second an AI model stops answering. We are drawing our core insights today from an exhaustive executive guide on operational resilience, specifically focusing on topic 8.8, which is titled continuity when the model goes down. This is really about treating operational resilience not as, you know, an IT afterthought, but as a core business function.
Because treating it as purely an IT issue is exactly how organizations find themselves paralyzed. Yeah. I mean when we talk about AI integration today, we spend so much time on the acquisition, right, the prompting, the deployment.
The shiny new toys. Exactly, but the resilience, the actual plan for when the magic box suddenly closes, is often left entirely to chance. And to anchor this entire discussion, to show you exactly why hoping for the best is basically a catastrophic business strategy, we need to look at a highly specific real-world event.
Let's take you back to 11.45 Coordinated Universal Time on the 30th of July 2024. A distributed denial-of-service attack, commonly known as a DDoS attack, hit Microsoft's Azure front door and the Azure content delivery network. Let's actually define that for a second, just to be completely precise about the underlying mechanics here.
A content delivery network, or a CDN, is essentially a globally distributed network of servers. Its entire job is to cache content really close to where the users are, so things load quickly for you. Right.
And Azure front door is the traffic management layer sitting right in front of that. It is basically the gatekeeper for a massive share of the Internet's Microsoft-hosted applications. So we were talking about millions of businesses relying on this exact gatekeeper.
Now a DDoS attack is essentially a malicious traffic jam, right? Right, yeah. Bad actors flood the gatekeeper with fake requests, and they're just trying to overwhelm its capacity so legitimate requests can't get through. Now, under normal circumstances, an attack like this would have been an ordinary contained incident.
Cloud providers face these constantly. Oh, all the time. They have massive, highly automated defense systems built specifically to absorb and deflect this exact kind of garbage traffic.
But that is where the architecture of this specific incident becomes, honestly, a master class in unintended consequences. If you read Microsoft's own post-incident account, there is a very sobering explanation of what actually happened. It wasn't the malicious traffic jam itself that caused the global outage.
What's fascinating here is that the malicious traffic triggered Microsoft's automated defense mechanisms exactly as designed. But an error in the implementation of those defenses caused the system to amplify the impact of the attack rather than mitigating it. Okay, let's unpack this.
Let's break down the mechanics of that. How exactly does a defense mechanism amplify an attack? Because on the surface, that sounds completely contradictory. It comes down to how these massive systems talk to each other under extreme stress.
So, when the defense mechanism kicked in to filter out the bad traffic, it inadvertently triggered a localized configuration error. Okay. And this error caused the system to start constantly querying its own internal infrastructure just to verify traffic rules.
Basically, instead of just dropping the bad packets of data and moving on, the defense system got stuck in this endless retry loop. It was essentially interrogating itself. Yes.
Asking itself over and over, millions of times a second, is this traffic okay? What about this one? What about this? So, the defense system itself became the flood. It swung with the punch. Precisely.
It swung with the punch. The internal routing was so overwhelmed by its own defensive queries that legitimate customer traffic was just dropped. The infrastructure essentially strangled itself trying to fight off the attack.
Wow. And the result of that self-strangulation was absolutely staggering. For roughly eight hours until about 19.43 UTC that evening, critical services like the Azure portal, Microsoft 365, and Microsoft Purview degraded or they just failed outright.
And they failed for customers who had done absolutely nothing wrong. Right. These businesses had no idea that a defense mechanism and not the initial attack was the reason their tools had suddenly gone dark.
And I think that brings us to the human element of that Tuesday morning. Because nobody who woke up that day had planned for this specific cascade of failures. Somewhere behind Azure front door that morning sat countless AI-dependent processes that organizations had never explicitly mapped as being vulnerable to a cloud outage.
You're talking about the invisible plumbing of the modern office. Exactly. I mean, customer service assistants answering billing questions in real time, triage tools routing complex support tickets, drafting tools that compliance teams leaned on every single morning without a second thought.
When Azure went down, all those tools went dark. Because gone, gone. Gone.
And for every single one of those localized business critical processes, exactly one question had to be answered that day. Whether anyone had bothered to ask it in advance or not. What happens now? Which brings us to the core premise of our deep dive today.
And I'm going to state this verbatim from our source material because it is the absolute spine of this discussion. Quote, every AI-dependent process has a default continuity plan already, whether anyone wrote it down or not. That is the hardest truth for executives to swallow.
You already have a plan, you just might not like what it is. Right. It's like, it's like a spare tire on a road trip, right? If you don't check if the jack works before you leave, your quote-unquote plan defaults to standing on the side of the highway waiting for a tow truck.
That is a great analogy because if nobody planned for an outage, if nobody deliberately designed a fallback, the default continuity plan is simply silence. Or it is a frantic expensive scramble by whoever happens to be on shift when the vendor's status page turns red. Yeah, an improvised queue of frustrated customers who have no idea why the usual instant answer has stopped coming.
And think about the operational friction that generates. Getting this wrong isn't just a technical embarrassment for the IT department. That is a massive misconception.
This is a business decision and if you haven't planned for it, it is a business decision made by default at the worst possible time, often by a mid-level engineer who entirely lacks the authority to actually make it. Right, because an unplanned outage on an unmapped AI process produces unbudgeted costs. Customers abandon their shopping carts, colleagues rack up over time trying to manually patch the gap, and eventually regulators or board members start asking why nobody was at the wheel.
Which is exactly why we are framing this as executive education today. By the end of this deep dive, you will have the exact framework to replace that default of silence and scrambling with an airtight tested plan. We are going to replace hoping the server stays up with an executable procedure.
So where do we start? Once we recognize that we desperately need a plan, how do we actually triage the dozens or maybe hundreds of AI tools embedded in a company's systems? Because we can't just judge them by how much money we spent on them or, you know, how much the CEO loves demonstrating them in board meetings. No, you have to completely mechanicalize the triage process. You have to remove ego and internal politics from the equation entirely.
And this brings us to the vital sorting metric. Again, going verbatim here. Classify on the harm of a gap, not on the process's importance.
Okay, let's unpack this because that is counterintuitive for a lot of leaders. Classify on the harm of a gap, not the importance. It is the single gate every AI dependent process must pass through.
When you are auditing your systems, you do not ask how important is this tool to our daily revenue. You ask one specific, very pointed question. If the model stops answering right now, what is the shortest amount of time before someone outside the organization is harmed, misled, or meaningfully inconvenienced? Okay, so we are measuring the velocity of the damage.
Exactly, the velocity of the damage. And based on the answer to that question, we sort every single AI process into one of three distinct lanes. Let's walk through them.
The first lane is what we call pausable. Pausable. So for a pausable process, the answer to our question about harm is measured in hours or maybe even days.
Correct. Nothing irreversible happens while the process waits. Think of a monthly report drafting assistant, an internal market research summarizer, or a marketing copy generator.
None of these produce a harm that grows exponentially the longer the pause lasts, provided, and this is key, provided that the pause is bounded and clearly communicated. Right, the business does not collapse if a marketing blog post goes out on Thursday instead of Wednesday. Exactly.
A pausable process can simply stop, hold its place, and wait for the model to return. All right, lane two, human fallback. In this lane, the timeline for harm is much tighter.
It is measured in minutes to maybe a few hours. A complete gap is unacceptable, but a person doing the work manually, even if they were doing it at a slower pace and at higher operational cost, is an acceptable substitute for that specific window of time. So examples here would be a customer support triage assistant, right, or maybe a claims intake summarizer for an insurance firm, or a first draft compliance flagging tool.
Exactly. Trained staff can physically step in and do the work while the model is down. The underlying business function doesn't grind to a halt, it just degrades in its efficiency.
The human basically becomes the bridge. And then we have the third lane, which is the highest stakes lane, cannot stop. Yes.
For a cannot stop process, the timeline for irreversible harm is measured in seconds, or it's a process where absolutely no human substitute exists at the speed or the volume the business requires to function. A gap of any real length produces harm that simply cannot be recovered, rolled back, or reversed afterward. Give me the classic example of this.
A real-time fraud screening model sitting directly in front of a live payment flow, or an AI-driven safety interlock model in an industrial operational system. You cannot simply wait for a human being to manually check these things because the volume of transactions is too massive, the latency requirement is too tight, we're talking milliseconds here, or the harm itself does not reverse. Right.
If a fraudulent transaction clears because the model is down, the money is gone. You can't undo it. I want to push back on this classification system for a second, just acting as a proxy for an IT director who might be listening.
What about a bespoke strategic planning assistant that is used exclusively by the C-suite? Say the CEO uses it every morning to synthesize market trends. That tool is incredibly expensive, it's highly visible, and it is considered extremely important to the leadership of the company. Why wouldn't that automatically be classified as cannot stop just based on the sheer organizational weight of the people using it? That is the most common trap in this entire continuity framework, and it is a trap driven entirely by internal corporate politics.
We have to ruthlessly separate importance from urgency. Yes, the C-suite's strategic planning assistant is highly important to the long-term vision of the company, but you have to apply the core question. Does anything time-critical break if it goes down for 12 hours? Well, no.
I mean, the CEO might be annoyed, but the company doesn't lose a wire transfer. Exactly. The C-suite can wait a day for a planning draft without breaking the business.
Therefore, it is a pausable process. The classification follows the harm of a gap, not how much the business values the tool on a daily basis. If you classify the CEO's drafting tool as cannot stop, you are allocating massive continuity budget to protect something that doesn't actually need protecting at that level.
Which means you are stealing budget from the systems that do. Precisely. That distinction is crucial.
Importance does not equal urgency. What are some of the other traps people fall into during this classification phase? Because I imagine the sheer volume of data can really muddy the waters. That's a second major trap, actually.
Confusing volume with urgency. Right. I can hear the marketing director right now saying, our AI generates 10,000 localized ad variants a day.
The volume is massive, therefore it must be cannot stop. Yes, and it's a completely flawed argument. You might have a system producing enormous volume, but you have to ask, if the model goes down, is any of that volume actually time sensitive on the scale of an outage? The answer is almost always no.
It can just pause. 10,000 paused marketing drafts still equals a pausable process. The sheer quantity of the output doesn't magically upgrade the urgency of the output.
Okay, here's another one. What about the model itself? Say an organization standardizes on a single major model. Can the IT department just say, all right, we use GPT-4 or CLAWD-3 for these 15 different tasks across the company.
We've assessed the model and we classify it as a human fallback class risk. Absolutely not. And that is the third trap.
Classifying the model abstractly instead of classifying the specific process it enables. Unpack that for me. Why is that dangerous? Because an AI model is just an engine.
You don't classify the engine, you classify the vehicle it is currently powering. The exact same foundational model can power a pausable internal HR memo tool on one server and it cannot stop real-time customer-facing translation tool on another. The classification belongs to the downstream consequence and the people it touches.
It never belongs to the model in the abstract. So the exact same vendor, even the exact same foundational model, can and should appear on your continuity table multiple times with completely different classifications attached to it depending on what it's doing. That's exactly right.
And the final trap to avoid when you're building this table is assuming symmetry across the people the process touches. What do you mean by symmetry? Let's say you have an AI driven search tool. If it goes down for internal staff, they might be perfectly patient.
They know there's an outage, they go grab a coffee, they work on something else for a bit. From their perspective, the process is pausable. But if that exact same search tool is quietly feeling a live public-facing e-commerce page with a promised instant load time, the customer is not going to be patient.
No, they are going to close the tab and go to a competitor. Exactly. You always have to classify the process against the least patient, most exposed party it actually touches.
All right, so let's say we've done the hard work, we've mechanically sorted our systems, we haven't let ego or volume or the vendor name confuse us. We have a spreadsheet that clearly identifies what is pausable, what is human fallback, and what is cannot stop. We're feeling good.
Which is usually where companies stop, unfortunately. Right, because a label isn't an action. Knowing a process is human fallback doesn't actually solve the problem at 11 for 5 a.m. when the Azure server crashes.
We need an executable procedure. Which brings us to our next verbatim spine point. A degraded mode is a specific procedure with a named owner, never a vague intention.
Yes, this is where disaster recovery plans usually devolve into complete fiction. Writing down the operations team will handle it in a spreadsheet is a hope, it is not a plan. A degraded mode must name, in plain language, exactly what happens instead of the AI output and exactly who or what is responsible for executing it.
It has to be incredibly clear. It has to be written so clearly that a competent stranger with 30 seconds to read it at 2 a.m. could execute the fallback without having to ask a single clarifying question. Let's break that down deeply by our three classes, because designing these modes requires very different thinking for each lane.
Starting with pausable. If the plan is just to let it pause, what is there to actually design? Doesn't it just sit there? For a pausable process, the degraded mode requires a visible cue, an honest message. It is never just silence.
If the system stops accepting new work, or if it holds requests in a visible cue, every single person interacting with it must see a specific pre-written message. That message must state exactly what is paused, roughly how long it is expected to remain paused, and where to go if the wait is completely unacceptable for their specific situation. Because a pausable process with no visible message isn't pausing gracefully.
To the end user, it is just a broken system failing silently. They don't know it's paused, they think they did something wrong, so they hit refresh 10 times, which just sends more garbage traffic to your broken system. Exactly.
And the most critical part of this degraded mode is that you must name who drafts and approves that message before an outage ever happens. Let's talk about the mechanics of that. Why does the owner need to be named in advance? I mean, can't someone just write a quick tweet? Because if you have a cue and message design built into the software, but no named owner authorized to deploy the wording, you will end up with an improvised, inconsistent scramble.
Marketing will want to say one thing to protect the brand. Legal will panic and say another to avoid liability. IT will just want to post an error code.
And while they're arguing on a conference call, the customer is staring at a blank screen. Exactly. The plan must say, in the event of an outage, the shift lead is authorized to deploy pre-approved message template B. Done.
Executable. Okay, moving to the human fallback lane. We've established we can't just say operations will take over.
What does a specific procedure look like here? For human fallback, the degraded mode needs three highly specific engineered elements. First, a specific named team or role. Not a department, a role.
Second, a specific manual procedure they follow, which is usually a slower, strictly defined manual version of what the automation was doing. And third, and this is where almost everyone fails, it requires a calculated volume threshold. The volume threshold.
Let's really dive into the mechanics of this because I think this is the silent killer of continuity plans. Let's look at the queuing theory behind it. Let's say your AI automatically routes 500 complex customer claims a day.
It reads the claim, identifies the core issue, and sends it to the right department. The AI goes down. Your human operations team takes over.
But you have to ask the hard math question. How many claims can that human team actually read and route manually in an hour before they are completely buried? Right, because a human reading a complex claim might take five minutes. The AI took five milliseconds.
Exactly. So if you run the math and the answer is that the team can handle 50 claims an hour, then your human fallback plan technically breaks at claim number 51. The queue starts backing up, wait times explode, service level agreements are breached.
You're just moving the failure point. Yes. You have to calculate that exact threshold where the human team gets overwhelmed.
And the degraded mode must name the escalation procedure for when that specific number is hit. Do they start bulk routing claims to a general queue? Do they trigger an overflow call center? Assuming a human fallback team can absorb AI level surge volume infinitely is a recipe for operational collapse exactly when you need your people the most. That makes perfect sense.
You're quantifying the breaking point of your own people. And now the highest stakes lane, the cannot stop class. We already know you can't use humans here.
The velocity of the harm is too fast. So what does a degraded mode actually look like in this lane? Cannot stop, the degraded mode is always a second technical path. It cannot rely on human intervention.
It could be a simpler rules-based system running in parallel that takes over automatically. Yeah. It could be a cache last known good decision provided it's applied very conservatively.
Or it could be a completely redundant AI provider that is entirely separate from the primary one. And this brings up a crucial concept from the source material regarding cannot stop processes. The failsafe default.
What exactly is a failsafe default? Cannot stop fallbacks must explicitly state their operational posture in plain language. If the primary AI fails and the fallback technical path kicks in, say it's a very rigid conservative rules-based system, what is its default posture when it encounters something it doesn't understand? Right, because an AI might be able to parse nuance but a rigid rules-based backup won't. Precisely.
So you must define the default. Does it deny the transaction? Does it automatically escalate it to a high priority supervisor queue? Or does it halt the system entirely? This cannot be left implicit. The continuity table must explicitly read deny, escalate, or halt.
Because in the heat of the moment you don't want ambiguity. Exactly. A colleague or a systems engineer executing this plan under extreme pressure at 2 a.m. needs to know the exact default posture to take without having to wake up a vice president to seek permission.
Let's transition here because this brings up a fascinating tension. We have the procedure. We've named the owner.
We have the volume thresholds. And we have the failsafe defaults. But how fast does this all need to happen? And who actually gets to decide that speed? Because I guarantee you if you ask a lead infrastructure engineer how long it should take to failover, they will give you a highly technical answer based on server spin-up times and database replication.
Which violates our next verbatim spine point. Recovery objectives belong to the business, not the technical team alone. Yes, and this is a fundamental misunderstanding of what operational resilience actually is.
Technical teams do not own the tolerance for the business harm that occurs on the other side of a delay. A recovery objective invented by a technical team in isolation is a number that simply won't survive the first real test. Walk me through the dynamic of that.
Why does an IT driven number fail in reality? Because it's based on technical convenience, not business reality. Let's say IT determines that failing over to a backup system will take 45 minutes because that's how long the data sync takes. They write that down as the recovery time.
But the process is a live customer checkout flow. Oh wow. The very first time an actual outage happens and real customers start abandoning carts by the thousands, a business leader is going to kick open the door of the operations center and demand the system be brought online immediately.
They will overrule the technical teams invented 45-minute timeline mid-crisis, causing absolute chaos. So to prevent that chaos we need two essential numbers and they must be fiercely negotiated and co-signed by the business side of the house. Number one is the maximum tolerable outage, the MTO.
This aligns with what standard IT calls the recovery time objective or RTO. Yes. The MTO is the longest amount of time the process may be down, running in its degraded mode before the financial or reputational harm to the business becomes entirely unacceptable.
For a plausible process, the business might say the MTO is four days. For human fallback, it's usually measured in hours, specifically bounded by that volume threshold we calculated, how long the staff can sustain the manual work before burning out. And for cannot stop, the MTO is measured in minutes or less because the fallback technical path needs to be live almost continuously.
And the second number, which is arguably even more in cities if you get it wrong, is the maximum acceptable staleness, the MAS. This aligns with the recovery point objective or RPO. This isn't about time offline, this is about data age.
How old can the last AI output be before it is deemed unreliable? Precisely. Let's say your cannot stop fallback is designed to rely on a cash fraud score. The primary AI goes down, so the fallback system just uses the last risks or it calculated for that user an hour ago.
Let me jump in here with an analogy because this is where the mechanics of data staleness get really dangerous. Serving a stale cashed fraud score is like serving yesterday's weather report to a pilot who is taking off today. That's a perfect way to look at it.
If that pilot relies on a 24-hour old wind shear reading, the results are catastrophic. The data technically exists, the system is technically up and providing an answer to the pilot, but the reality that data describes has completely changed. A storm has moved in.
A cash decision with no explicit expiry date quietly becomes a confidently wrong answer. And in many ways, a confidently wrong answer from a machine is far more dangerous than an honest visible system failure. That is exactly the risk of ignoring the maximum acceptable staleness.
If a model stopped updating an hour ago, that output might still be perfectly safe for a low stakes internal drafting task. But for a fraud risk score, getting a live thousand dollar payment, the behavioral reality that score was computed against has already changed in that hour. Naming the MAS in advance prevents a technical team from quietly trusting an output that is silently gone stale just because nobody thought to set an expiry date on the cash.
Setting and signing these two numbers, the maximum tolerable outage and the maximum acceptable staleness, in a calm conference room with the business relationship owner is the entire value of this exercise. You write the numbers down and you write the name of the vice president or director who agreed to them. Because a recovery objective with no named business co-owner is functionally the same as an objective the technical team invented alone, this signature turns the metric into a business decision that someone can actually be asked to defend during an audit.
Which moves us into the reality check. So we have the plan. The business signed off on the numbers, the spreadsheet is beautiful, everyone feels great.
But as we saw with the Azure front door incident at the top of the show, a beautiful spreadsheet means nothing when the servers actually melt. Microsoft designed a defense mechanism. On paper it was supposed to mitigate DDoS attacks.
In reality, under hostile load it created a retry loop and strangled its own network. So our next verbatim spine point, a described fallback and a tested fallback are different things. This distinction is everything in operational resilience.
A paper test is just corporate theater. It is reviewing a document in a meeting. Everyone sits around a conference table, the risk manager reads the graded mode procedure aloud, everyone nods and says, yes this makes sense.
But that catches nothing. That kind of test catches absolutely nothing. It won't catch the retry loops, it won't catch the stale cache, it catches nothing that a real outage won't catch later at a far higher price.
So what does a genuine test look like? Walk me through the mechanics of proving this works. A genuine failover test requires you to actually pull the plug. You trigger the degraded mode in a controlled environment.
You measure the actual activation time of the fallback system against that signed maximum tolerable outage. If the MTO is two minutes and the fallback takes five minutes to spin up, you fail. You actually stress test the humans too, right? Absolutely.
You put simulated volume through the human fallback team to confirm that their capacity actually holds up to the threshold you calculated. You confirm that the downstream customer actually sees the intended, cleanly written degradation message and not a terrifying string of 502 bad gateway code errors. So if testing is this rigorous, should we just test everything uniformly? Like put it on a master calendar, every single AI process in the company gets a genuine test once a year to keep it simple and fair across all departments.
No, and that is a massive, incredibly common misconception that drains IT budgets. A uniform calendar misallocates resources. Testing every process on the exact same schedule wastes immense effort on the low-stakes rows.
But much more dangerously, it creates a false sense of security. How so? It gives leadership the illusion that your high-stakes rows have been rigorously validated just because a checkmark appears in a spreadsheet once a year. Ah, I see.
The test cadence has to be sized to the classification lane. Exactly. A cannot-stop process like our real-time fraud model deserves a frequent, disciplined testing cadence, monthly or even tighter.
The cost of discovering a broken fallback during a real cannot-stop outage is exactly the irreversible financial harm that class exists to prevent in the first place. Meanwhile, a plausible process, like the monthly report drafter, can be tested semi-annually or even annually. The cost of a delayed discovery there is comparatively small.
You focus your testing budget where the harm is fastest. That brings up a terrifying realization about underlying architecture. Let's say you've done everything right.
You've classified the process as cannot-stop. You've built a redundant technical path. You test it genuinely every single month.
But what if your primary AI model and your fallback redundant technical path go offline at the exact same time? This is our next verbatim spine point. Different vendor names can still share one underlying point of failure. We call this concentration risk, and it is the boogeyman of modern cloud architecture.
Let's say your organization uses five different, highly reputable, vendor branded AI tools. You have one for customer service, one for drafting, one for identity management, and so on. You look at your continuity table, you see five different logos, and you feel heavily diversified.
Sure, you think you've spread the risk around. But if all five of those separate vendors run their back-end infrastructure on the exact same cloud region, say AWS on EastOne, or they all rely on the exact same content delivery network, you haven't diversified your risk at all. You've just hidden it behind different marketing logos.
It's infrastructure over branding. You can't just look at the software label, you have to look at the metal it runs on. One upstream failure, like the Azure front door routing error, takes out all five of your quote-unquote diversified vendors simultaneously.
And I want to be clear that this is no longer just a theoretical best practice or a thought exercise for IT nerds. For regulated financial entities, this is now the law. The European Union's Digital Operational Resilience Act, known as DORA, became fully binding on the 17th of January 2025.
Oh wow, so it's a legal requirement now. Yes. DORA legally mandates that financial entities map and test exactly this specific ITT third-party concentration risk.
Yeah. Regulators are demanding to know if your primary and secondary systems share the same hidden dependencies. Now obviously, if we are living in reality, you can't engineer away every single shared dependency.
There are only so many major cloud providers in the world. Trying to build completely independent stacks for every tool would basically bankrupt most IT departments. True.
You cannot eradicate concentration risk entirely, but you must make it visible. You must price it as an exposure. Two processes with different vendor names, but the same underlying cloud region, are just one single point of failure wearing two different labels.
If you discover this shared dependency during your mapping, you name it explicitly on a continuity table. You force the business to accept the risk consciously, or you force them to fund a compensating control. What you never do is let it sit undiscovered, simply because nobody bothered to ask the vendors how their infrastructure works under the hood.
Let's make all of this theory incredibly concrete. Let's look at how this entire framework, the classes, the limits, the testing, the budgets, applies in the real world when you are constrained by a fixed budget. We're gonna introduce a character to anchor this.
Velma. She is the operations risk lead at a fictional midsize insurance company called Harbor Line Mutual. Velma is the one sitting at her desk when Azure goes down.
She has meticulously mapped out five specific AI dependent processes in her organization. Let's walk through her continuity table. Alright, process one at Harbor Line Mutual is an internal policy drafting system.
If it goes down, the drafting team just waits. It's classified as plausible with a maximum tolerable outage measured in days. Process two is a legal document summarization tool used for discovery review.
Same thing. The court deadlines are usually weeks out. It is plausible, measured in days.
Process three is a customer facing coverage chat bot on the main website. Now this is highly visible. Customers interact with it directly to ask about their deductibles, but if it goes down, the widget is designed to gracefully degrade.
It just shows a pre-written message with a 1-800 phone number. So Velma classifies this as plausible as well. Now it gets interesting.
Process four is a back-end claims triage system. When a customer submits a claim, the AI reads it and routes it to the correct adjuster. If it goes down, claims still come in, they just aren't routed automatically.
Velma classifies this as human fallback. And she has to do the math here. Yes.
She sits down with the VP of claims and they agree on a maximum tolerable outage of four hours. But Velma also calculates the volume threshold. Harbor Line receives roughly 400 new claims a day.
Velma calculates that her three trained triage staff members can manually read and route about 90 claims in that four-hour window. Beyond that 90, a massive backlog builds and service level agreements are breached. She documents that exact limitation.
And finally, process five. A real-time fraud flagging model sitting right in front of live claims payouts. If this model stops answering, payouts either halt entirely, which the business vehemently refuses to accept because it violates state insurance regulations, or they proceed blindly without fraud checks, which is financially catastrophic.
This is definitively cannot stop. Right. Velma secures a maximum tolerable outage of two minutes and a maximum acceptable staleness of zero.
You absolutely cannot serve stale fraud score on a real payout. So Velma has the table. But this raises the ultimate defining question of operational resilience.
The budget. Velma has a fixed budget. Let's say she has limited engineering weeks available this quarter to actually build and test these degraded modes.
How does she allocate that limited resource? Well, she funds the cannot stop fraud flagging model first. It requires a completely redundant, separately hosted, rules-based path that is incredibly expensive to build, and it requires disciplined monthly testing. And in our scenario, a business director pushes back on Velma.
He looks at the spreadsheet and asks why the highly visible customer chat bot on the home page only gets tested twice a year, while the back-end fraud model, which customers never even see, gets the lion's share of the budget and gets tested monthly. And Velma's answer is the defining logic of the continuity budget. She tells the director, customers seeing a clear pause message on the chat bot is a bad afternoon.
A fraud check silently going dark on a live payout is a bad year. Budget follows consequence, not visibility. Exactly.
You spend your engineering budget on the row where a gap cannot be undone, even if it's the most expensive to build, and you fund the cheap, low-consequence, pausable rows last. If Velma had done what most companies do, an even split of the budget across all five processes, to be fair, she would have produced a shallow, half-tested, degraded mode on every row. And a half-tested fallback on a cannot stop fraud model is functionally identical to having no fallback at all.
When the Azure outage hit, Harborline would have hemorrhaged cash. But here is where the beautiful spreadsheet meets reality. Even a business-aligned, ruthlessly prioritized plan like Velma's will hit snags when reality gets messy.
Let's cover the edge cases, the things that break the plan. Starting with partial degradation, or what I like to call the slow leak. The Azure front door incident is actually the perfect example of the slow leak.
It wasn't a clean binary break. It wasn't just on at 11.44 and then completely off at 11.45. It was intermittent errors. It was random timeouts.
It was severe cascading latency spikes where queries took 10 seconds instead of 10 milliseconds. That kind of erratic behavior is genuinely difficult for automated monitoring to distinguish from just an ordinary bad network day until it's already been ruining your downstream systems for an hour. Walk me through the mechanics of that.
If your continuity plan is built only around a clean binary, meaning the AI model either answers or it doesn't, you are going to completely miss the partial degradation. Your plan will never trigger. Exactly.
The system will just bleed out slowly. Your monitoring must be engineered to trigger degraded modes based on specific granular signal thresholds. You don't just monitor for a total outage.
You track latency spikes over a five-minute rolling window. You track rising refusal rates from the API. You track a sudden inexplicable collapse in the model's confidence scores.
So you're looking for symptoms of failure, not just a flatline. Yes. You define the specific threshold for those metrics, and when it's crossed, you trigger the degraded mode automatically.
You cannot sit around on the Slack channel waiting for the failure to become undeniable to everyone downstream. Edge case number two. Surge volume and false positives.
We talked about Velma's human fallback team handling 90 claims during their four-hour window, but outages inherently cause demand surges? They absolutely do, and it's a brutal psychological cycle for the staff. If a customer-facing system goes down, what do customers naturally do? They call in to complain about it. Of course.
So your human fallback team is suddenly hit with their normal manual routing work, plus an enormous unexpected surge in angry complaints about the outage itself. Fallbacks must be explicitly sized for surge volume, not just baseline normal volume. And how do false positives play into this, the crying wolf scenario? This is the flip side of the monitoring thresholds we just talked about.
If your monitoring threshold is set too sensitively, trying to catch every slow leak, it will trigger the degraded mode during ordinary network jitter. That false positive failover has a massive real-world cost. How bad is it? It creates unnecessary manual work for Velma's team.
It sends alarming degradation messages to customers when nothing is actually wrong. And worst of all, it erodes the operations team's trust in the alert system itself. The next time the alert goes off, and it's a real outage, the team will react slowly because they assume it's just another false alarm.
Monitoring thresholds must be tightly mathematically calibrated. Okay, the final edge case, the reconciliation gap. The Azure outage is finally over.
It's 19.43 UTC. The status page is green again. The AI model is back online.
So we're done, right? High fives all around the operations center. Not even close. The work is just beginning.
What happened to all of those 90 claims that were routed manually by Velma's team over the last four hours? What happened to the thousands of transactions that were conservatively denied by the rules-based fallback system? A continuity plan that just says the model is back, return to normal, is wildly incomplete. Right, because you have two different data realities now. You have what the AI knows, and you have what the backup systems did while the AI was blind.
Exactly. A proper continuity plan names exactly how the degraded mode's output gets reconciled back into normal operations. Those manually routed claims need to be uniquely tagged during the outage so they can be audited later, or merged seamlessly back into the primary system without causing data corruption or duplication.
If you don't plan for the reconciliation gap, the outage just creates a permanent scar in your data architecture. We have covered incredible ground today. We are moving into our final takeaways.
We've built the table, we've classified the urgency based on harm, we've designed specific degraded modes with named owners, we've forced the business to set the recovery numbers, we've scheduled genuine tests, we've audited for shared infrastructure, and we've planned for the messy edge cases like the slow leak. But there is one final rule. A continuity table is a living artifact.
Yes. A continuity table goes stale the exact moment your underlying business processes change. It is not a static compliance document you write once, print out, and file in a drawer just to show an auditor.
It must be reopened on specific operational triggers. What kind of triggers? At contract renewal for any AI vendor, at the first material change to any of the covered processes, or at minimum on a fixed inflexible review calendar. Look at Harborline Mutual.
If the business suddenly starts promising customers a same-day turnaround on a summarization tool that used to take a wake, that tool is no longer plausible. It is now human fallback, and the table must be immediately updated to reflect the new velocity of harm. So to you the professional listening to this, what is the single most valuable move you can make this coming Monday morning? It is remarkably simple.
Start with the systems inventory you already have. Do not try to boil the ocean. Pick just five AI-dependent processes, just five.
Block out exactly 30 minutes with the business owner of those processes. Right. Classify them into the three lanes.
Write out the specific degraded mode. Get the maximum tolerable outage and the maximum acceptable staleness signed by that business owner. And schedule a genuine failover test.
Turn the default of silence into a deliberate executable plan. And as we close this deep dive, I want to leave you with a final thought to mull over. It builds on the human fallback class we've discussed so heavily today.
A continuity plan that relies on a human fallback is quietly making a massive unstated claim about your workforce capacity and cross training. It implies roles that don't just use AI, but actually act as its shadow architecture. The shadow architecture, the people catching the falling glass.
Exactly. Ask yourself this. Does your HR department, do the people actively mapping your workforce capacity and hiring targets actually know that their people are the foundational load-bearing walls of your IT continuity plan? Because if they don't, when the model goes down and you reach for that fallback plan, well you're gonna find that the load-bearing wall simply isn't there.
Real cases
Example 1 (the anchor): the Microsoft Azure Front Door outage, 30 July 2024. A distributed denial-of-service attack against Azure Front Door and Azure Content Delivery Network produced an unexpected usage spike; Microsoft's own account states plainly that an error in the implementation of its defenses amplified the impact of the attack rather than mitigating it. The resulting disruption ran roughly eight hours, from 11:45 to 19:43 UTC, and degraded or broke the Azure Portal along with a range of Microsoft 365 and Purview services. Read backward: the failure was not the attack, which cloud providers face constantly and generally absorb; it was a defense mechanism behaving differently under real, hostile load than it was assumed to behave, precisely the gap a genuinely exercised failover test, not a described one, is built to catch. (Sources: Microsoft Azure status history and post-incident summary, 30 to 31 July 2024; Help Net Security, "Microsoft: DDoS defense error amplified attack on Azure, leading to outage," 31 July 2024; Reuters and Dark Reading contemporaneous coverage.)
Example 2 (supporting, pointer): insurers excluding AI liabilities. The deep treatment of the shrinking insurance market for AI-related loss belongs to Topic 8.4 (see Topic 8.4). This topic's addition: an uncovered continuity gap is not a hypothetical exposure once several major insurers have begun filing to exclude AI-chatbot and agent liabilities from standard policy language; a process with a weak or absent degraded mode is, in effect, a self-insured tail the organization is carrying without having decided to, the same framing Topic 8.4 applies to a thin liability cap.
Example 3 (supporting, pointer): the NHS Federated Data Platform's exit and hosting terms. The deep treatment of this contract's clauses belongs to Topic 8.7 (see Topic 8.7). This topic's addition: exit-management provisions and hosting-location restrictions negotiated at signing are a form of long-horizon continuity planning, protecting what an organization can still do if a vendor relationship ends badly, which is the same discipline this topic applies to a much shorter horizon, what an organization can still do in the hours a vendor is simply unavailable.
Example 4 (illustrative pattern, not a single named incident): the shared-dependency surprise. A recurring, documented pattern across cloud-outage postmortems industry-wide (used here as an illustrative structure grounded in ordinary cloud-architecture practice, not a single named AI case): an organization believes it has diversified risk across several vendor-branded tools, discovers during an outage that every one of them sits on the same underlying cloud region or content delivery network, and experiences a full, simultaneous outage of tools it had assumed were independent. Read backward: a continuity table built only at the vendor-name layer, without checking the infrastructure layer beneath it as Section 3F describes, can look diversified and still fail as a single point of failure the first time that shared layer goes down.
Example 5 (supporting, pointer): the two a.m. incident and the value of a plan already written. The deep treatment of running a live AI incident response belongs to Topic 3.5 (see Topic 3.5). This topic's addition: every improvised decision that a responder has to make live during that two a.m. call (whether to pause, who takes over manually, how stale the last output can be trusted) is a decision a completed continuity table has already made in advance, in a calm room, with the business owner's agreement, turning the incident response from an invention exercise into an execution of a known plan.
Example 6 (illustrative pattern, not a single named incident): the tested fallback that was never actually tested. A pattern documented repeatedly across operational-resilience postmortems in banking, telecommunications, and cloud services (used here as an illustrative structure, not a single named AI case): an organization's continuity documentation describes a fallback in confident, specific language, a named team, a named procedure, a named recovery time, and the fallback is discovered, only during a real outage, to have quietly stopped working months earlier (a phone tree with disconnected numbers, a script referencing a system that was decommissioned, a team that was reorganized without anyone updating the plan). Read backward: a written degraded mode is a claim, not a fact, until a genuine, scheduled test exercises it and confirms the claim still holds.
Example 7 (supporting, pointer): model drift as a quieter cousin of an outage. The deep treatment of quiet model degradation belongs to Topic 4.5 (see Topic 4.5). This topic's addition: the maximum-acceptable-staleness discipline built here for a cached fallback output applies equally to a model that never technically went down but has quietly drifted past the point its outputs should be trusted without re-validation, showing that a sudden outage and a slow drift are two different causes converging on the same governance answer, a defined expiry on trust.
Where people go wrong
- "Continuity planning is only for cannot-stop processes." Every process needs a named degraded mode, including the pausable ones. A pausable process with no queue-and-message design does not fail gracefully; it fails silently, which is a worse customer experience than a clearly communicated pause, even though the underlying harm is small.
- "Importance and continuity class are the same judgment." A strategically important process (a quarterly planning assistant leadership relies on) can still be pausable, because nothing time-critical happens while it waits. Classify on the harm of a gap, not on how much the business values the process day to day.
- "A described fallback is a tested fallback." A document stating what should happen during an outage proves nothing about what will happen. Only a genuine, scheduled exercise of the degraded mode, with the fallback actually triggered and measured, closes the gap between the plan on paper and the plan's real behavior, the same gap the Azure Front Door defense mechanism fell into.
- "The technical team can set the recovery objective alone." A recovery number invented without the business owner's sign-off will not survive the first real incident; someone with more authority will overrule it on the spot, and the plan will have been worthless exactly when it mattered.
- "Different AI vendor names mean diversified risk." Two processes running on different AI products can still share the same underlying cloud region or content delivery network. A continuity table that checks only the vendor layer, not the infrastructure layer beneath it, can look diversified and still be a single point of failure.
- "A human fallback works at any volume." A human-fallback plan sized only for normal volume can collapse exactly when it is needed most, since an outage often coincides with a surge in the very demand the fallback exists to absorb. Name the volume threshold honestly and name what happens beyond it.
- "A cached last-known-good answer is always safer than no answer." A cached decision with no staleness expiry can quietly become a confidently wrong answer rather than an honest gap. Set the maximum acceptable staleness in advance and fail safe (deny, escalate, halt) once it is exceeded, rather than serving an aging cache indefinitely.
- "Testing every process on the same schedule is thorough." It is actually a misallocation. A uniform test calendar spends the same effort on a low-stakes pausable row as on a cannot-stop row, and creates a false sense that both have been tested with equal rigor when only one of them needed to be.
- "A continuity table, once built, does not need to change." A table built for the AI-dependent processes an organization runs today goes stale the moment those processes change, exactly the way a model-change-notice clause protects against a vendor's model changing underneath the contract (see Topic 8.7). Reopen the table at renewal, or the first time any of its five processes materially changes, whichever comes first.
- "The vendor's incident-notification clause is the continuity plan." A strong notification clause tells the organization sooner that a problem exists; it does not, by itself, tell anyone what to do about it. The clause and the continuity table answer two different questions, and winning the first does not excuse skipping the second.
- "Building the redundant path for a cannot-stop process is a one-time engineering cost." A redundant technical path that is built once and never exercised again drifts out of sync with the primary path it is meant to replace, as data formats change, as thresholds are retuned, or as the primary system's behavior evolves. The fallback needs the same ongoing maintenance and the same scheduled testing as the primary path, not a single build-and-forget project.
- "A continuity table is a compliance document, not an operational one." A table built only to satisfy an audit checklist, without a genuine intention that a colleague could execute it during a real outage, fails at the one moment it exists for. Write the table to be used at two in the morning by someone who has never seen it before, not merely to be filed after a review.
- "Once the outage ends, the degraded mode's job is done." Whatever the degraded mode produced while it was active, a queue of paused requests, a batch of manually routed claims, a set of conservatively denied transactions, still has to be reconciled against normal operation once the primary path returns. A continuity plan that ends at "the model is back" without naming how the transition back is handled leaves a gap of its own, the reconciliation gap named in Section 3G's fifth edge case.
Questions people ask
- What is continuity table?
- The artifact this topic produces: a row per AI-dependent process, each carrying a continuity class, a degraded mode, two recovery numbers, and a last genuine test date, defensible to a hostile board member line by line.
- What is continuity class?
- One of three categories (Pausable, Human-fallback, Cannot-stop) a process is sorted into based on the shortest time before someone outside the organization is harmed, misled, or meaningfully inconvenienced by the model's silence.
- What is pausable process?
- A process where nothing time-critical happens while the model is unavailable; the process simply queues or waits, with a visible message, and resumes when the model returns.
- What is human-fallback process?
- A process where a trained person can perform the model's function manually, at a slower pace and usually a higher cost, for a bounded window before the fallback team itself becomes overwhelmed.
- What is cannot-stop process?
- A process where no human substitute exists at the required speed or volume, and any real gap produces an irreversible harm, requiring a redundant technical path rather than a pause or a person.
Keep going
This lesson builds AI resilience, continuity and exit planning, and that page shows the roles that hire for it. Every Certified AI Governance Professional (CAIGP) lesson.