Business leaders should require an exit plan before an AI model or platform becomes essential to customer service, finance, security, operations, software delivery, or another critical workflow. The plan must show what the business owns, what the provider controls, how required data can be recovered, which replacement has already passed evaluation, who can authorize a migration, and how the minimum service continues while technology changes underneath it.
An AI exit plan is not a promise that every model is interchangeable or that a multi-provider library can fail over automatically. It is a tested operating capability: preserve the business process and control evidence, restore an acceptable level of service within approved recovery targets, and keep a manual path when the replacement cannot safely reproduce the preferred system. Build it before an outage, deprecation, export restriction, contract change, regional limit, or provider access decision forces the schedule.
Why August 20 changed the resilience conversation
An August 20 TechRadar Pro opinion asks the question many AI programs defer: what happens when a capability embedded in business operations suddenly becomes unavailable? Its core distinction is useful. Security attempts to protect a system from compromise; resilience determines whether the business can continue when a system, service, or data source is unavailable, regardless of cause.
The same date provided a concrete lifecycle example. Google Cloud's Grok 4.1 Fast model notice says the xai/grok-4.1-fast-reasoning and xai/grok-4.1-fast-non-reasoning endpoints were deprecated on the Gemini Enterprise Agent Platform and would be shut down on August 20, 2026. After that date, Google's Agent Platform Model as a Service would no longer serve those endpoints. Google directed customers to newer xAI models such as Grok 4.2 or Grok 4.3, or another model in Model Garden.
That notice is about Google Cloud's serving surface. It does not say xAI shut down Grok 4.1 everywhere. The distinction matters because an AI dependency has layers: the model developer, cloud or inference host, product platform, region, API version, identity system, quota, contract, safety service, retrieval store, tool connection, and business workflow. A model can still exist while one route to it disappears. An exit plan must identify the route the company actually depends on.
Google's general open-model deprecation guidance defines deprecation as a notice period during which an existing endpoint may still work but receive no new features and restrict new use. Retirement permanently deactivates the endpoint, causing calls to the retired model ID to fail. That difference should drive two separate dates in the change calendar: the migration decision deadline and the last safe production-use date.
Vendor dependency is a workflow risk, not a model preference
A model name is only one row in the dependency register. A production AI workflow may rely on a provider's SDK, streaming format, structured-output behavior, tool-call schema, content filters, regional endpoints, identity exchange, caching, vector storage, evaluation service, observability, batch mode, file API, assistants or agents runtime, and contract terms. Replacing the text-generation call while those dependencies remain unresolved is not a failover.
Provider-specific behavior can also be valuable. A workflow may perform well precisely because a model follows a certain prompt, supports a particular context length, calls tools in a predictable way, or generates a schema that downstream software accepts. Portability does not require suppressing those advantages. It requires making them visible, isolating the provider-specific portion, and deciding what degraded but safe service is acceptable when exact parity is unavailable.
NIST's AI Risk Management Framework Core supports that operating approach. It calls for inventories based on risk priority, safe decommissioning, documented third-party and supply-chain risk, contingency processes for failures in high-risk third-party AI systems, viable non-AI alternatives, and assigned authority to disengage systems whose outcomes no longer match their intended use. An exit plan turns those governance outcomes into an executable runbook.
The nine-control AI vendor exit plan
Use this matrix as the acceptance contract for every critical AI workflow. Replace generic roles with named people or on-call groups, attach the evidence to the workflow inventory, and set a review date. A diagram, vendor questionnaire, or unused backup account does not pass the control until the team demonstrates the stated exit test.
| Exit control | Accountable owner | Required evidence | Pass condition |
|---|---|---|---|
| 1. Dependency register | AI platform owner and workflow owner | Model IDs, endpoints, regions, SDKs, identities, quotas, tools, data stores, contracts, and named owners | Every production dependency maps to an owner and an approved replacement or manual path |
| 2. Criticality and recovery | Business continuity owner | Business impact tier, maximum outage, RTO, RPO, minimum service, and restoration priority | Recovery targets reflect business impact and are approved before production use |
| 3. Data portability | Data owner, privacy, and legal | Export formats, retention and deletion terms, log and memory treatment, timing, cost, and tested archive | Required records can be exported, read, reconciled, and deleted on the agreed clock |
| 4. Portable workflow boundary | Application architecture owner | Versioned prompts, schemas, tool contracts, business rules, provider adapters, and routing configuration | A provider change does not require rebuilding the business process from a closed workspace |
| 5. Evaluated replacements | Model risk and product owner | Approved candidates, representative test set, current evaluation date, limits, and degraded-mode decision | At least one replacement meets the minimum quality and control threshold for each critical workflow |
| 6. Production qualification | Security, platform, finance, and quality owners | Authentication, quota, region, latency, schema, safety, observability, cost, and load-test results | Replacement capacity and controls work under realistic demand—not only in a demo |
| 7. Manual operating path | Business operations owner | Manual intake, queue, prioritization, approval, reconciliation, backlog, staffing, and return-to-service steps | People can sustain the minimum service for the maximum planned manual window |
| 8. Trigger and authority | Executive risk owner and incident commander | Migration triggers, decision thresholds, contact tree, change authority, rollback authority, and communications | The on-call team can name who decides, switch, stop, notify, and reverse without an ad hoc meeting |
| 9. Failover exercise | Resilience program owner | Scenario, timeline, RTO and RPO results, accepted-output rate, control parity, cost, backlog, and corrective actions | A scheduled exercise meets targets and every missed target has an owner and due date |
1. Inventory the complete dependency chain and its owners
Start with the business outcome and trace every external and internal component required to complete it. Record the workflow owner, business users, model developer, serving platform, exact model and API version, endpoint and region, SDK or gateway, orchestration layer, system and developer prompts, retrieval index, embeddings model, files and memory, tool integrations, identity and secret owners, safety filters, output schema, monitoring, evaluation suite, billing account, quotas, contract, support path, and downstream consumers.
Record both visible and transitive dependencies. A customer-support assistant might call one branded model but rely on a separate cloud host, identity provider, search API, vector database, ticketing connector, content-moderation endpoint, and tracing vendor. Ask what fails if each layer disappears, changes terms, rejects the account, loses a region, exhausts quota, or stops returning the expected schema.
Give every dependency one operational owner and one business owner. The operational owner maintains configuration and recovery instructions. The business owner decides the minimum acceptable service and impact of downtime. Procurement, security, privacy, data, and legal owners should be linked where they control contracts, evidence, or approval. Unowned dependencies are not low risk; they are unmanaged risk.
2. Classify workflows by business criticality
Classify the workflow by impact rather than token volume or executive enthusiasm. Consider customer harm, revenue delay, safety, financial reporting, legal or regulatory duties, employee access, operational backlog, recoverability, data sensitivity, and the point at which a manual queue becomes unmanageable. One high-volume marketing workflow may tolerate days offline while a low-volume fraud, incident, or production decision cannot.
For each tier, set the maximum tolerable outage, recovery time objective, recovery point objective, minimum service level, restoration priority, and maximum period of manual operation. RTO is the target time to restore the minimum service. RPO is the maximum acceptable loss of state or work since the last recoverable point. NIST's contingency-planning guidance connects recovery plans to business impact and explicitly includes alternate technology and short-term manual processing.
Do not assign an aggressive RTO without funding the required capacity, evaluated replacement, credentials, on-call coverage, data synchronization, and rehearsal. A written 15-minute target with an unopened replacement account is not a recovery capability. Approve the cost and service tradeoff at the same level that accepts the business impact.
3. Document data export, retention, and deletion before exit
List every record the company must preserve or move: prompts, prompt templates, uploaded files, vector-store source documents, embeddings or their reproducible inputs, conversation state, agent memory, tool definitions, fine-tuning files and outputs, evaluation sets, traces, approval records, generated artifacts, feedback, audit logs, billing records, and configuration history. For each, record the system of record, owner, export method, format, encryption, retention period, deletion process, restoration test, transfer time, egress cost, and contractual right.
Do not assume the provider's general privacy statement describes every feature. Google's Vertex AI retention guidance, for example, documents different behavior for abuse monitoring, search or maps grounding, session resumption, and in-memory caching. The lesson is broader than one cloud: verify retention and training terms for the exact service, feature, agreement, configuration, and geography used by the workflow.
Exportability and usability are different tests. A JSON archive passes only when an independent process can read it, reconnect identifiers, rebuild the necessary state, reconcile record counts and hashes, and satisfy retention or deletion duties. Test large exports against the recovery clock. Keep a protected, provider-independent copy of the prompts, schemas, source documents, evaluation cases, and configuration that the company is entitled and required to retain.
4. Separate prompts, tools, and business logic from provider APIs
Keep the workflow definition in a versioned system the company controls. Store business rules, prompt templates, tool contracts, input and output schemas, routing policy, approval requirements, evaluation cases, and migration configuration outside a provider-only console when the product permits it. Expose provider-specific calls through narrow adapters so the business process does not directly depend on one SDK throughout the codebase.
This boundary is not a universal lowest-common-denominator wrapper. It should describe the business contract—such as classify an intake, return evidence, draft a response, request approval, or propose a transaction—then allow each adapter to use the target provider safely. Record which capabilities have no substitute, which require prompt changes, which tools need a different schema, and which outputs must be downgraded to human review.
Keep secrets, identities, quotas, regional configuration, logging, safety policies, and retry rules outside the prompt. Version each deployment so responders can answer which model, adapter, prompt, tool registry, evaluation release, and policy handled a particular transaction. Portability without provenance can move the workload while destroying the audit trail.
5. Maintain evaluated replacement models
Build a replacement bench for the workflow, not a generic list of popular models. At least one candidate should be available through a dependency path that does not share every failure mode with the primary. Two model names served by the same cloud endpoint, identity tenant, region, gateway, and billing account may improve model choice without providing provider-level resilience.
Evaluate candidates on a fixed set of representative and adversarial cases. Include common work, edge cases, ambiguous inputs, long context, multilingual or multimodal cases if relevant, expected refusals, prompt injection, restricted data, tool errors, stale retrieval, conflicting evidence, schema violations, and known incidents. Define the minimum accepted-output rate, correction rate, safety behavior, latency, reviewer effort, and cost before seeing the candidate's results.
Re-run the suite when the primary or replacement model, prompt, safety settings, tool schema, retrieval source, SDK, or routing layer changes. A candidate evaluated six months ago against an earlier workflow is an assumption, not a ready replacement. Keep the evaluation date, test-set version, limitations, approver, and permitted degraded mode with the dependency record.
6. Qualify the replacement as a production system
A model-quality score does not prove production readiness. Test authentication and least-privilege roles, credential rotation, endpoint and regional availability, network controls, contract activation, quotas and rate limits, reserved or provisioned capacity where needed, concurrency, context and file limits, streaming, timeouts, retries, idempotency, structured outputs, tool calls, safety filters, logs, alerts, support escalation, and rollback.
Load-test the replacement at realistic peak demand and under a cold start. Verify quota increases before the exercise rather than assuming the provider will grant them during a major outage affecting many customers. Test whether the secondary account, region, or provider shares identity, DNS, network, payment, or cloud-control-plane dependencies with the primary.
Measure cost per accepted business task, not only price per million tokens. A nominally cheaper model can cost more if it uses additional output tokens, increases retries, requires more review, breaks tool calls, or produces work that must be corrected. Capture output quality, elapsed time, reviewer minutes, safety exceptions, failed transactions, and total provider and tool charges for the same evaluation workload.
7. Preserve a manual operating path
Some incidents will make automation unsafe before they make it impossible. A provider may return degraded output, lose a safety control, fail to preserve required evidence, or change terms in a way that needs review. The continuity plan must let an authorized person move the workflow into a visible manual queue without continuing uncertain AI calls.
Document intake, prioritization, staffing, access, approved templates, evidence sources, dual control, exception handling, customer communications, backlog limits, reconciliation, and return-to-service. Identify which work stops, which work proceeds manually, and which service level can be promised. Prebuild the queue and permissions; a PDF that tells employees to use a system they cannot access will fail during the incident.
The manual path should preserve transaction IDs and decisions so completed work is not repeated when automation returns. Reconcile items generated, queued, approved, sent, posted, or paid across the interruption boundary. If the business cannot sustain the manual path for the maximum outage, lower the allowed dependency, add capacity, or fund a stronger technical alternative.
8. Define migration triggers and decision authority
Write objective triggers before commercial or operational pressure clouds the decision. Examples include a confirmed outage exceeding the failover threshold, a deprecation notice inside the validated migration lead time, repeated quality or safety regression, loss of an approved region or data term, a quota reduction, an export or sanctions restriction, a material price or contract change, a security incident, loss of required support, or a provider decision that removes account or model access.
For each trigger, name the observer, evidence source, decision owner, technical executor, business approver, security and privacy reviewers, communications owner, and rollback authority. Define when the on-call incident commander can invoke the tested degraded mode immediately and when an executive risk owner must approve a longer migration. Include vendor contacts but do not make a vendor response a prerequisite for protecting the business.
The runbook should state whether traffic will stop, drain, queue, route by workflow tier, or switch in full; how in-flight work is reconciled; what users and customers are told; what safety and approval restrictions tighten in degraded mode; and what evidence allows the primary provider to return. A kill switch, failover switch, and return-to-service gate are separate controls.
9. Run failover exercises with measurable recovery targets
Exercise one critical workflow at a time in a safe environment, then progress to a controlled production slice when risk owners approve it. Do not tell the operating team which dependency will fail. Inject a provider outage, revoked credential, retired model ID, quota reduction, contract hold, unavailable export, or output-quality regression and require the team to diagnose the layer rather than blindly change the model name.
Measure detection time, decision time, time to minimum service, RTO and RPO attainment, successful export and state recovery, backlog growth and clearance, accepted-output rate, correction rate, safety and approval parity, failed tool calls, quota headroom, cost per accepted task, manual capacity, communications time, and return-to-service accuracy. Record every missed target with one owner, due date, verification test, and risk decision.
An illustrative exercise might interrupt an AI-assisted customer-support triage workflow. The team must preserve incoming cases, route high-severity cases to people, authenticate to the approved replacement, apply its adapter and safety policy, run a canary set, confirm evidence links and ticket schemas, cap volume, monitor correction rate, clear the manual backlog, and reconcile every case. This is a planning example, not an ITECS client incident.
Run exercises on a risk-based schedule and after a material change to the primary provider, replacement, workflow, identity, retrieval source, tool, contract, or recovery target. Tabletop discussion can validate roles; a technical exercise must prove credentials, capacity, data, adapters, controls, and outputs. Alternate scenarios so the team does not memorize one scripted answer.
A 30-day path to an exit-ready pilot
During the first week, select one recurring critical workflow and complete the dependency and owner register. Classify its impact, set minimum service, RTO, RPO, maximum manual window, and migration authority. Open every relevant contract and data term; do not infer exit rights from product marketing.
During the second week, export and restore required records, move workflow definitions into version control where allowed, isolate provider calls behind an owned boundary, and prepare the manual queue. Choose one replacement whose infrastructure path removes the failure mode being tested.
During the third week, run the representative evaluation and production qualification: identity, network, quota, load, region, safety, schema, tools, observability, latency, quality, and total cost. Document the degraded mode and stop if the replacement misses a minimum control; a rushed migration should not trade availability risk for unsafe output.
During the fourth week, conduct a timed failover exercise, reconcile all work, verify return to service, and present the missed targets to the risk owner. Repeat until the workflow meets its approved targets. Then apply the same contract to the next-highest criticality workflow rather than declaring the entire AI portfolio resilient from one successful test.
Google Cloud's Grok 4.1 notice made one endpoint transition visible, but a provider does not need to disappear for a critical workflow to stop. Models retire, APIs change, regions and accounts lose access, quotas bind, contracts move, features retain data differently, and output quality can fall below the business threshold. Leaders who inventory those dependencies and rehearse the exit can make a deliberate migration. Everyone else discovers the architecture during the outage.
