← Ideas & guides

Guide

When AI becomes unavailable: prepare for continuity and vendor exit

A guide for leaders to define minimum operations, identify dependencies and rehearse recovery using a fictional case. Includes preparation for a provider change.

By Dr. Sven JungmannRetrospective reference date: · Published: · Reviewed:
Interlocking index cards form a supporting bridge.

At a glance

  • Decide which tasks must continue during an interruption.
  • Two applications can depend on the same underlying service.
  • A fallback becomes credible when responsible people have rehearsed it using permitted test data.

A useful AI application can quickly become part of a workflow. At some point, a new question arises: what happens to that workflow if the application is unavailable tomorrow or needs to be replaced permanently? The answer concerns procurement, professional leadership, information technology and the people who actually do the work.

This guide proposes a bounded joint exercise. It clarifies minimum operations, important dependencies and an orderly return to normal working. It is an organisational aid. Existing clinical, technical and operational emergency procedures must be considered separately by the people responsible for them.

Start with the result that remains necessary

Consider a fictional example. A pharmaceutical company uses an AI application to assemble approved internal documents for a weekly leadership meeting. The service fails on the morning before the meeting. Knowing who can contact the provider is insufficient on its own. Leaders need to decide which information must still be available and which work can wait.

The first joint question is therefore: what result must remain available during the interruption? In this example, a short list of already confirmed decisions might be sufficient. New analyses can move to the next meeting. This agreement limits the scope of the fallback and makes the accepted reduction in service explicit.

Since 2010, NIST SP 800-34 has described a planning approach that connects interruption impacts, required resources and recovery priorities. It addresses US federal information systems. I use its mechanisms as a professional reference; they do not establish a German legal obligation for the fictional company. [1]

Make four dependencies visible

Draw the workflow on one page: incoming information, processing, checking and handover. Add what each step depends on. Four categories are sufficient for an initial discussion: required information, technical services, authorised people and binding working rules.

In the example, original documents may be held in a separate approved repository. The AI application also contains the sequence in which they are assembled. An experienced employee knows the exceptions. Even this short description reveals which elements remain accessible during an outage and which would need to be reconstructed.

Examine shared technical dependencies as well. Two different applications can use the same underlying AI service, authentication system or data repository. Provider information and the organisation's architecture can establish whether this is the case. A second interface alone does not demonstrate an independent fallback.

The exercise does not require a publicly accessible list of internal systems. Its findings belong in the appropriate protected documentation. Access credentials, keys and confidential information must remain within approved procedures during an interruption too.

Test the fallback through a concrete task

For the fictional meeting workflow, an assigned person could gather confirmed decisions directly from the approved repository. A second person checks the time period and selection. Recipients are told which additional analyses are missing on this occasion. This describes a limited workflow whose requirements can be examined.

Ask a named deputy to carry it out using artificial documents. They should be able to find the required information, understand the sequence and reach the intended handover. Record where a question becomes necessary. This can reveal dependence on knowledge that has never been made explicit.

The duration of the exercise is an observation under those conditions. It cannot serve as a guaranteed recovery time for every outage. An actual incident may simultaneously affect staffing, access or additional systems. Leaders should therefore state the type of interruption for which the procedure is intended.

The voluntary NIST generative AI profile of 2024 recommends considering third-party dependencies, fallback arrangements, responsibilities and incident exercises. It is a risk-management framework, rather than evidence that the exercise proposed here is effective. [2]

Recovery includes reconciling unfinished work

When the service works again, the organisational task remains unfinished. In the example, decisions were assembled manually during the outage. The application could later process the same documents again. Without reconciliation, duplicate or contradictory versions may result.

Agree in advance who authorises the transition back. That person needs to identify work already completed, results still missing and intermediate versions that can be discarded. A short list recording status and responsibility may be sufficient if it suits the scale of the activity.

Recipients also need to know when the ordinary workflow resumes. A technical success notification does not automatically answer that organisational question. The accountable leader decides whether the necessary checks have been completed and processing can restart.

Prepare separately for a permanent provider change

A longer transition requires more than a temporary workaround. It involves information that can be transferred, documented working rules, rights to the content used and conditions of future use. These questions should be addressed early with procurement, the accountable subject owner and the relevant legal and privacy functions.

Request a sample of what can be exported in a usable format for the workflow in question. Check whether professionals can reconstruct the meaning of the work from it. A collection of old conversations may be incomplete: it could contain many answers while lacking a reliable current rule. What is available in the actual product needs to be checked against the agreed service.

A replacement application also requires its own professional assessment. Identical inputs and similar-sounding outputs do not guarantee equivalent processing. Previously approved example tasks can help prepare a comparison; responsibility for the future use must be assigned again.

A planned change also needs an explicit handover point. The organisation should know which system is authoritative for each piece of work during the transition. Temporary coexistence may be necessary, but people need a clear rule for resolving competing versions and directing new work.

The decision after the exercise

The exercise should leave a small number of clear results: accepted minimum operations, responsibility for activating them, a rehearsed fallback, conditions for recovery and unresolved dependencies. Every open issue should have an owner and a next step.

Leaders can then deliberately decide which additional precautions justify their cost. For me, this is the exercise's value: it translates a vague concern about dependence into concrete decisions about how the organisation can continue to act when conditions change.

Sources and further reading

  1. NIST SP 800-34: impact, recovery and exercisesNIST

    The 2010 guide concerns recovery of US federal information systems. Its planning mechanisms provide a reference here; it does not establish a German legal obligation.

  2. NIST: generative AI value-chain risksNIST

    The voluntary 2024 profile recommends fallback procedures, ownership and exercises for failures of third-party AI systems. It does not guarantee the safety of a particular architecture.

Perspective and interests

This article was developed with AI assistance. The organisational examples are fictional. The practical proposals are original inferences from the sources within their stated limits.

I am the founder and CEO of aiomics and have a commercial interest in the responsible adoption of AI in medicine.

Keep reading

What should your event make possible?

Tell me about your audience, occasion and timing. We can shape a talk around the questions that matter to them.

Enquire about a talk