The AI Nobody Put on the Continuity Plan
- By Winston Thomas
- October 05, 2026

Pick a Monday. Gather your leadership team and set one rule: Your primary AI vendor is gone until next Monday. No API, no model and no status page promising a fix.
The first casualty is obvious: customer service. At organizations leading in AI, customer-facing assistants resolve up to 70% of customer inquiries without a human, says Kitman Cheung, technical sales director for ASEAN at IBM Technology. Switch them off, and “an AI outage will have a serious impact on a company’s ability to serve its customers,” he says.
The second casualty is worse: the list of everything else that stops. “I think most CDOs will find that AI usage is much wider than anyone anticipated,” Cheung says. An agent might run one step in pricing and procurement, and another might check documents during client onboarding. Each small job can take down the systems that depend on it, he says.
IBM’s data backs up the worry. In the IBM Institute for Business Value study “The calculus of AI sovereignty,” 81% of 1,000 executives said a seven-day outage at their primary AI vendor would be severe or critical. Only 9% said they had an excellent understanding of their AI dependencies. Most companies know the fall will hurt, but only a few have measured the height.
The drill is no longer just theoretical. On June 12, a U.S. export-control order led Anthropic to suspend access to two of its frontier models for all customers. But Cheung’s warning is less about Washington than about the dependencies already inside your own company.
Minimum viable sovereignty
Finding those AI dependencies starts with a definition. Cheung divides digital sovereignty into four parts: operational (who runs the environment), data (who controls it), technology (whether your architecture can leave a vendor) and AI (where models run and how inference, the work a model does when it responds, is governed).
AI is the hardest of the four, he says, because hardware limits and vendor lock-in cap how much control any organization can hold. “Many analysts agree that it is impossible to be fully independent in 2026,” he says.
So full independence is not the goal: A minimum level of control is. “In the AI era, organizations need to assess their own level of minimum viable sovereignty to balance cost and risk,” Cheung says.
Think of it as the minimum viable product of control, a framing Forrester analysts popularized: the least control you need over a system to survive the week it fails. The IBM study calls its version selective AI sovereignty: Spend on control where failure costs the most, and accept lock-in, knowingly, everywhere else.
But you can’t set that minimum until you know what you depend on.
Under the floorboards
Security teams learned how costly hidden dependencies are in December 2021. Google found that most vulnerable copies of Log4j, a widely used open-source logging library, sat more than one level deep in the software stack, with some sitting nine levels down. The U.S. Department of Homeland Security estimated it would take at least a decade to find and fix them all.
Cheung’s list of hidden AI dependencies reads like the sequel.
Start with data. Organizations inherit shadow training data: datasets with unclear origins. “Datasets used to train models are often invisible, yet models’ behaviors are nearly completely dictated by training data,” he says. The cloud adds another blind spot. “Once it is in ‘the cloud,’ where the AI model actually caches and processes the data is often a vendor black box.” For a CDO accountable for data residency, that is a hole in the map.
“I think most CDOs will find that AI usage is much wider than anyone anticipated.” — Kitman Cheung, IBM Technology
The model layer is more challenging. A developer connects a third-party API to a data pipeline to process images, summarize documents or extract customer intent. The project then ships, and the developer moves on. But the API call keeps running. Those forgotten calls, Cheung says, “can propagate and stay invisible.” Then there are model-on-model chains, where one model’s output feeds another model. Even if the first model is mapped, he warns, the chain behind it can stay hidden. Software engineers call this a transitive dependency. It is the Log4j problem again, with an inference bill.
Infrastructure adds two more risks, Cheung says. AI compute is concentrated in a few regions, which exposes organizations to trade restrictions and capacity shortages. And the chip supply chain depends on a handful of global suppliers.
Regulators are starting to ask for this map. The Monetary Authority of Singapore’s proposed AI risk management guidelines would require financial institutions to identify and inventory all AI systems they use. The draft covers AI agents, too.
Yet in a stress test, Cheung expects the first failure to come from people and process, not technology. “I suspect it will be the operating model that breaks first,” he says. “Generative AI arrived with a bang.” Companies rushed into pilots, and some pilots moved into production without the controls production requires. “They move into production without SLAs, operational runbooks or monitoring,” he says. No one owns validation, model drift or outage response.
His fix is simple. Tie every project to a business goal with measurable KPIs. Name an owner for each model, training dataset and set of test cases, and an operational owner for deployment, monitoring and recovery.
Portability dies in the last mile
Owners fix accountability. Portability, the ability to move data, swap models and shift workloads, fixes lock-in. It is also where the industry has been burned before.
Service-oriented architecture, virtualization and early cloud all promised portability. The promise still outruns reality: 75% of executives who switched, or tried to switch, AI vendors in the past two years told IBM the process was difficult.
Cheung points to shortcuts at the end of implementation as a common culprit. A proprietary connector got an enterprise service bus into production faster. A hypervisor was portable on paper, but production ran on a vendor-specific control plane. The architecture diagram said open, the runbook said otherwise.
AI has its own last mile, and it is longer than the API call. A model swap usually means re-evaluating outputs and retuning prompts. In banking, the IBM study notes, a model change can also trigger revalidation and audit cycles. If the swap includes the embedding model, which turns text into numbers for search, every vector in the retrieval store must be regenerated, because embeddings from different models are not compatible. Even IBM’s own Granite embedding models come in 384- and 768-dimension versions. Switch from one to the other, and the old index no longer works with new queries.
Leading CDOs engineer that last mile on purpose, Cheung says. They place a model gateway, or routing layer, between the application and the model, so the application never calls a vendor directly. They run governance, monitoring and traceability in a layer above the models, so swapping a model does not erase the audit trail. They keep data in open formats with full lineage, so a new model can train on trusted data without rebuilding the pipeline. And they design for failover. “The same architecture that allows vendor swaps can also help an organization survive an outage,” Cheung says.
Some executives already draw that line. “We accept long-term strategic relationships at the application layer,” Dalton Gouws, group IT director and board member at Volkswagen Group UK, says in the IBM study. “But, at the model layer, we deliberately preserve optionality.” In short: Commit to the application vendor, but keep the model replaceable.
IBM learned the same lesson on its own estate. Its approach “took shape as an evolutionary journey,” Cheung says. Platform sprawl taught the company that portability must be engineered, not promised. Lock-in from earlier technology waves pushed it toward open source and Red Hat OpenShift. IBM research also found that organizations applying hybrid-by-design principles can earn more than three times the return on investment over five years.
The confused deputy
So far, the risks involve AI that answers questions. Agents go further by taking actions. That changes the sovereignty question from where data sits to who authorized what.
IBM’s Responsible Technology Board lists four agent traits that introduce risk: opaqueness, open-endedness, complexity and non-reversibility. Open-endedness is the easiest to underestimate. Agents choose their own tools, resources and even other agents, so there is no fixed code path to audit.
That opens the door to a classic security flaw, the confused deputy problem: A program with legitimate privileges is tricked into using them for someone else. An agent with broad credentials is a deputy waiting to be confused.
“The same architecture that allows vendor swaps can also help an organization survive an outage.” — Kitman Cheung, IBM Technology
Cheung starts with identity. “Agents are digital workers that participate in work,” he says. “They need to have their own identity with sufficient role-based access control.”
An agent that runs on a shared service account leaves a log showing what happened, but not on whose behalf. The OAuth token exchange standard offers a cleaner pattern by naming the user as the subject and the agent as the actor, so every call carries both identities. Multi-agent chains have a catch, though. Under the standard, only the current actor is enforceable, and earlier agents in the chain are recorded for information only. Five agents deep, you have a record, not five checkpoints.
This is not only a security issue. For a CDO, an agent’s action is a data lineage puzzle: who touched which data, and on whose authority? When the trail stops at a shared service account, the lineage graph has a gap.
That is why Cheung’s questions matter. Which tools can an agent use without human approval? Can you trace its actions across organizational boundaries? On whose behalf was it acting? And the question CISOs will ask first: “Which actions are reversible? Which are not and should be gated with human oversight?”
One model, three countries
These questions get harder when one AI system serves several countries. Cheung described a hypothetical bank, drawing on insights from IBM’s 2025 CDO Study and the IBM white paper “Digital sovereignty in the era of hybrid cloud and AI.”
The bank operates in Singapore, Indonesia and Malaysia. It wants one AI system for fraud and credit risk that serves all three markets. MAS expects strong AI governance. Indonesia requires certain customer data to stay in the country; its Financial Services Authority, OJK, requires banks to keep data centers onshore unless it approves otherwise. All three regulators want to audit AI decisions.
The CDO, Cheung says, can split the problem in four ways: Regulated data stays in-country. A data fabric, a layer that provides governed access to data where it sits, lets the model use that data without moving or copying it. The model trains on de-identified data using a regional hyperscaler’s compute. And a separate governance layer spans all three environments, with local policies for each regulator’s audits.
That governance layer carries real weight, because de-identification has limits. Research by privacy scholar Latanya Sweeney found that ZIP code, birth date and gender alone uniquely identify 87% of Americans. Remove names from a training set, and the remaining combinations can still point to a person. Residency is a data rule, and scale is a compute rule. Governance lets the two coexist.
The budget line
Every layer in that design costs money, and the IBM study says to plan for it. Its advice to CFOs is to fund optionality like insurance. That means a sovereignty premium based on how critical each system is, plus contracts with enforceable data portability rights and exit SLAs.
The return is measurable. Organizations with the strongest control across their AI stack protect 55% more operating profit from AI-driven disruption than organizations with less control. “Vendor lock-in creates imbalance,” says Conor Mlacak, chief information officer of Staples Canada, in the study. “Once you’re locked in, you lose leverage.”
That is the case a CDO can take into a budget meeting. The premium buys leverage, and the drill shows whether it is working.
Back to Monday
Cheung’s 12-month playbook asks the CDO to own four decisions. The first is the AI dependency map for key business processes. The second is the placement architecture: Keep regulated data sovereign, provide governed access without moving data and use regional hyperscalers for heavy compute where rules allow. The third is a governance layer that works across models and vendors. The fourth is the drill itself, along with guardrails for agents. The IBM study supplies the scorecard: time-to-switch, cost-to-switch and the share of critical systems with validated alternatives.
So run the drill. Write down everything that breaks, especially what nobody knew existed. Next quarter, run it again and check whether the list is shorter.
“Prove continuity,” Cheung says. “Don’t assume it.”
Image credit: iStockphoto/Zoran Jesic
Stay ahead with CDO Nexus
Join an exclusive community of CDOs and data leaders
- Access curated trends and thought leadership from IBM and industry experts.
- Participate in private roundtables and webinars to solve regional CDO challenges.
- Connect with a network of like-minded leaders navigating the same data landscape.
Winston Thomas
Winston Thomas is the editor-in-chief of CDOTrends. He likes to piece together the weird and wondering tech puzzle for readers and identify groundbreaking business models led by tech while waiting for the singularity.