Ch-Ch-Changes
Failed AI Pilots Are Tuition; AI Debt Is the Interest

Field Guide
Two threads pulled through one knot: what the “AI projects fail” statistics actually measure, and a name for the bill nobody itemizes — AI debt. Then the bridges organizations are building over a fast-moving ecosystem, ranked by the strength of the evidence.
There was a rhythm to enterprise software, and your organization knew it by heart.
Patch Tuesday. The quarterly feature release. The major version every eighteen to twenty-four months, announced at the vendor conference a year before it shipped, with a migration guide and a deprecation runway measured in years. The hardware refresh you could see coming three budgets away. None of it was fun, but all of it was danceable: you staffed an upgrade team, negotiated a test window, convened the change advisory board, and kept up. Keeping up was a solved problem. It was, at worst, a line item.
Then the rhythm section quit. In March 2023, OpenAI shut down its Codex API — the model a generation of coding tools was built on — with less than a week's notice. ChatGPT plugins launched that same month, spawned an ecosystem of third-party developers, and were fully retired thirteen months later. GPT-4.5 lived four and a half months in the API — shorter than most enterprise procurement cycles. The Assistants API, pitched at DevDay 2023 as the way to build agents, never left beta before its replacement shipped; anyone who architected around it owns a forced migration with a hard 2026 shutdown date. And in April 2026, OpenAI put more than twenty-five model IDs on notice for shutdown this year. This is one vendor. Add Anthropic, Google, Meta, DeepSeek, and the tool layer churning on top of all of them, and one tracker puts the median gap between frontier model releases — across labs — at about seventeen days in 2025.
Meanwhile, your organization has begun its AI adoption. Probably three initiatives, maybe five. And somewhere in a leadership deck is a statistic saying 95 percent of them will fail. This post pulls on both threads — the failure number and the pace problem — because they turn out to be the same knot.
- Frontier model release
- Deprecation / forced migration
- Paradigm shift
The old rhythm — any 3½ years of the SaaS era
The AI ecosystem — 2023 to mid-2026
Locked in early 2024, by mid-2026 the ground has moved
24 times
12 frontier releases · 6 deprecations or forced migrations · 6 paradigm shifts. In the old rhythm, the same span held about 10 quarterly updates and 1 major upgrade — with roadmaps that told you they were coming.
Frontier releases shown through 2025 — trackers disagree on 2026 model naming, and at the measured cadence (one tracker puts the median gap between frontier releases across labs at ~17 days in 2025) the dots stop being individually legible anyway. Which is the point. Sources in the notes below the post.
I
Autopsy of a Statistic
What the failure numbers actually measure
The number you've seen is 95 percent, and it comes from MIT's Project NANDA — “The GenAI Divide,” published in the summer of 2025 and made famous by a Fortune headline that briefly spooked the market. Read the fine print and the claim shrinks on contact: 52 interviews, 153 surveyed leaders, and a definition of failure — no measurable P&L impact within about six months of the pilot — that is doing heroic amounts of work. Wharton's Ethan Mollick put it plainly: the methodology doesn't support a general failure rate, and most pilots not scaling is normal for first experiments. Critics noted the 95 percent figure can't even be cleanly derived from the report's own tables. It is the most-quoted number in enterprise AI, and it is the methodologically weakest.
It also has company, and the company doesn't agree with itself. RAND's much-cited “more than 80 percent of AI projects fail” arrives with the words “by some estimates” attached — it's a borrowed, pre-GenAI figure; the actual RAND contribution is a root-cause list from 65 practitioner interviews, topped by leadership misunderstanding the problem to be solved. S&P Global found 42 percent of companies had abandoned most of their AI initiatives — self-reported, fielded in late 2024 and early 2025 — while the same survey found GenAI adoption rising. Gartner's famous numbers (30 percent of GenAI projects abandoned after proof-of-concept, 40 percent of agentic projects canceled by 2027, 60 percent doomed without AI-ready data) are predictions, not measurements; Gartner's one recent measured figure, from a survey of 782 infrastructure leaders published in April 2026, found about 20 percent of AI use cases failing outright — and the leading self-reported cause, at 57 percent, was “expecting too much, too fast.” Not the model. The expectation.
Here's the context the headlines never include: these numbers are ordinary. The Standish CHAOS report put general IT project success at 16 percent in 1994 and 29 percent two decades later, with large projects almost never succeeding. ERP implementations fail to meet objectives at somewhere between 55 and 75 percent — Gartner predicts about 70 percent of recent ERP initiatives will still be failing their business cases in 2027. CRM's canonical failure number ran 50 to 70 percent circa 2002. Big data projects: 60 percent, later revised by the same analyst to “closer to 85.” Digital transformation: 70. Every enterprise technology wave for thirty years has produced a canonical “most projects fail” statistic, every one of them was contested in its own time, and AI's sits comfortably inside the band. If your instinct is that AI projects fail anomalously often, the honest reading of the record is stranger: enterprise IT projects fail this often, always have, and we keep being surprised.
II
Tuition
What a dead pilot buys, and when it buys nothing
So: are failed AI projects really failures? The best evidence for “no” hides inside the pessimistic reports themselves. The same MIT study behind the 95 percent found that workers at more than 90 percent of surveyed companies were regularly using personal AI accounts for work — a “shadow AI economy” delivering results while the official pilots stalled. The technology was succeeding in the same buildings where the projects were failing. That's not a verdict on AI; it's a verdict on the wrapper — the workflow design, the integration, the org chart around the model. MIT called the root cause a “learning gap.” A learning gap is a thing experience closes, and pilots — including dead ones — are how the experience gets bought.
The named cases read the same way. McDonald's ended its hundred-restaurant AI drive-thru test with IBM in 2024 — a failure, by headline standards — while stating the test “has given us the confidence that a voice-ordering solution for drive-thru will be part of our restaurants' future,” and moving on to its next AI partnership. That is a company describing tuition, in so many words. Klarna replaced seven hundred support roles with an AI assistant, discovered that optimizing purely for cost “led to lower quality,” rehired humans — and kept the AI, rebuilt as a hybrid. The maximal deployment failed; the failure specified the working design. And the broader pattern has data behind it: an MIT Sloan Management Review and BCG study of more than 3,000 managers found organizations that build multiple modes of human-AI learning are roughly six times more likely to see significant financial benefit. What predicts payoff isn't first-attempt success. It's learning-mode intensity — which failed attempts feed. Deloitte's surveys even show organizations planning for this: most expect fewer than 30 percent of their AI experiments to reach scale, and they increase investment anyway. A high pilot kill rate isn't a scandal. In a portfolio, it's the design.
Now the concession that keeps this honest: tuition only counts if somebody writes down the lesson. A pilot that dies in a post-mortem — here's the workflow that didn't fit, here's the data that wasn't ready, here's the eval we should have had on day one — is R&D. A pilot that dies in silence, because the sponsor changed or the budget cycle turned or the whole thing was ritual to begin with, is just a receipt. The difference between the organizations collecting tuition and the organizations collecting receipts is not the failure rate. It's the write-down. And there is a second, quieter way a pilot's value evaporates — which brings us to the other thread.
III
AI Debt
Naming the bill nobody itemizes
Software engineering already has a term for the corners you cut on the way to shipping: technical debt. Machine learning already has its own version — Google researchers wrote the canonical paper in 2015, “Hidden Technical Debt in Machine Learning Systems,” about the entanglements and data dependencies inside an ML system. What's accumulating in enterprises right now is a third thing, and it deserves its own name. Call it AI debt: the distance between the direction your AI initiative locked in — models, tools, architecture, vendor bets — and the direction the ecosystem has moved since. Classic tech debt is the gap between your code and what your code should be. AI debt is the gap between your bearing and the frontier's. You accrue it without writing a line of code. You accrue it by standing still.
Look at the figure above and pick the half-year your current initiative froze its architecture. Everything to the right of that line is your accrued principal. Locked in early 2024 — a perfectly reasonable moment, GPT-4 was a year old, the stack looked stable — and by mid-2026 the ground has moved more than two dozen times: agents arrived, MCP went from nothing to industry standard in twelve months, the API you built on was deprecated, and the model you benchmarked against was shut down. None of that is your team's fault. All of it is your bill. The old rhythm would have delivered ten quarterly updates and one major upgrade in the same span, with roadmaps and runways. The ecosystem doesn't publish a roadmap. It publishes a surprise, roughly biweekly.
The interest on this debt shows up in ways organizations have already felt but not named. Prompt engineering was a $335,000 job posting in 2023 and an obsolete role by 2025 — the models absorbed the skill; whole team structures were built around a discipline that evaporated in twenty-four months. Framework churn has its own literature now: teams that adopted heavy abstraction layers in 2023 to hedge against vendor change spent 2024 and 2025 unwinding them, because function calling, structured outputs, memory, and tool orchestration kept getting absorbed into vendor primitives. The abstraction you adopt to protect against change becomes the change you have to unwind. Even the vendors carry it: OpenAI's own GPT-5 rollout yanked a beloved model, restored it days later under user revolt, and produced a public apology from its CEO. If the lab shipping the frontier can't cleanly manage its own upgrade path, your steering committee's eighteen-month AI roadmap never stood a chance.
And here is the multiplier that makes this generation of debt different in kind, not just degree: the agents. The same tools your initiative deploys to produce more work produce more debt, at the same ratio, on the same schedule. GitClear, analyzing hundreds of millions of changed lines of code across the AI-assistant era, found duplicated code blocks up eightfold during 2024; copy-pasted code exceeding refactored code for the first time in the dataset's history; refactoring itself collapsing from about 21 percent of changes in 2021 to under 4 percent by 2026; and the share of work touching code more than a year old — maintenance, the unglamorous verb that keeps systems alive — down by nearly three quarters. Agents multiply output. Multiplication doesn't care what it multiplies. An organization running five AI initiatives with agent-accelerated delivery is compounding classic tech debt inside each project while AI debt compounds around all of them — a debt spiral with two axes, which is not a sentence anyone's quarterly business review has a slide for.
IV
The Thrash
Why leadership keeps changing direction — and why your team feels it
If you work anywhere near an AI initiative, you know the feeling this section is about. The priorities changed again. The pilot you spent a quarter on is suddenly the wrong architecture. Leadership read something on a plane. A survey by Writer and Workplace Intelligence — 2,400 employees and executives, published this spring — put numbers on the feeling: 54 percent of the C-suite said AI adoption is tearing their company apart, up double digits from the year before. Seventy-five percent of executives admitted their company's AI strategy is “more for show” than real guidance. Fifty-five percent called their organization's AI usage a chaotic free-for-all. Mainstream business press spent 2025 coining “AI fatigue” for what the people inside these companies already knew.
The easy story is that leadership is fickle. The honest story is worse and kinder at the same time: the thrash is structural. When the ground genuinely moves every few weeks — when the median gap between frontier releases is measured in days, and a category like “agentic coding” goes from demo to deployment inside a year — a leadership team that never changed direction would be negligent, and a leadership team that always changes direction is indistinguishable from chaos. Both failure modes are responses to the same input. Meanwhile the people not on an AI initiative can't keep up by definition — following this ecosystem is a full-time job that nobody assigned to them — so every directional change arrives without context, and feels like whiplash instead of navigation. The gap between how fast the ecosystem moves and how fast an organization can absorb direction changes is AI debt, experienced as morale. Same distance, measured in people instead of architecture.
V
The Bridges
What actually works, ranked by the strength of the evidence
You cannot avoid AI debt by picking the right model, the right framework, or the right vendor, because every pick depreciates in six to eighteen months. The organizations managing this well aren't the ones predicting the future. They're the ones making their picks cheap to change. Five approaches, strongest evidence first.
1. Own your evals; rent your model. An evaluation harness — twenty to fifty real tasks from your actual workflows, with graded answers — is the one artifact in an AI stack that appreciates while everything else depreciates. Anthropic's own engineering guidance says the quiet part: teams without evals face weeks of manual testing per model upgrade; teams with them re-run the suite and upgrade in days. The practitioners who teach this — Hamel Husain, Eugene Yan — converge on the same line: nobody regrets investing in evals. And the market shows the payoff at scale: enterprise LLM spending swung massively between vendors in under twenty-four months — one lab's enterprise API share falling by nearly half while another's more than tripled — which means organizations that could re-evaluate and switch actually did, and captured the gains. The ones locked in watched. Your eval suite is your switching passport. It is also, not coincidentally, where a dead pilot's tuition gets deposited: the failure teaches you what to test for.
2. Keep the routing layer thin. A gateway that makes models swappable — the LiteLLM/OpenRouter/Portkey category, or the multi-model planes inside Bedrock and Vertex — is the direct architectural answer to a world where one vendor can deprecate twenty-five model IDs in a single notice. But the LangChain saga is the standing caution: heavy frameworks are themselves a depreciating bet, and teams that adopted thick abstractions spent longer unwinding them than the abstractions ever saved. The rule that survives: abstract the seam (which model, which provider), not the paradigm (how an agent thinks) — the paradigm is what keeps changing.
3. Buy more than you build, and centralize the watching. Menlo Ventures' enterprise data shows the market already voting: 76 percent of AI use cases were bought rather than built in 2025, up from roughly half the year before — and MIT's data, for what it's worth, found bought solutions reaching deployment at twice the rate of internal builds. The reason is AI debt itself: a vendor amortizes the cost of keeping up across every customer; your internal build carries it alone. The complement is organizational: McKinsey finds most gen-AI adopters running centrally led AI functions even where everything else is federated — because one team tracking the ecosystem is a cost; fifty teams each tracking it badly is a catastrophe. One caveat from the alpha argument: buy the commodity layers, but the workflows that make you defensible are exactly the ones worth owning — build there, buy everywhere else.
4. Run a portfolio with an intended kill rate — and a tuition ledger. Small bets, short cycles, explicit success criteria, and a scheduled decision date on which most bets die. That's not failure management; that's the operating model the abandonment statistics accidentally describe. The addition almost nobody makes: a written record of what each dead pilot taught — the workflow mismatch, the data gap, the eval that would have caught it. Ten dead pilots with post-mortems are an education. Ten dead pilots without them are the 95 percent.
5. Remember the 10-20-70 rule. BCG's heuristic for AI transformation puts 10 percent of the effort in algorithms and models, 20 percent in technology and data, and 70 percent in people and process. If that ratio is even directionally right, then the layer churning fastest — the model layer — is the smallest part of your investment, and AI debt is survivable by construction: the 70 percent (how your people work, how decisions route, how output gets reviewed and directed) transfers across every model swap. Organizations drowning in AI debt are usually the ones that inverted the pyramid — bet the initiative on a specific model's specific behavior, then discovered the model was a rental. Defer the commitments that lock you to the churning layer until the last responsible moment; spend early on the layers that move slowly. The slow layers are where the actual savings live anyway.
VI
Turn and Face the Strange
The two threads, tied
Here's the knot untied. The failure statistics and the pace problem are the same phenomenon observed at two altitudes. Projects “fail” at the historically normal rate for enterprise technology — but each failure is booked as a total loss because the ground moves before the lesson can be reinvested. The organizations that feel broken aren't failing more than their predecessors did with ERP or CRM. They're failing at the same rate while the tuition depreciates faster — a lesson learned against last year's stack pays out only if your switching costs are low enough to apply it to this year's. Which is why every bridge in Section V is really the same bridge: evals, thin routing, buying the commodity, killing pilots on schedule, weighting the slow layers — each one is a way of lowering the cost of changing your mind.
The old rhythm is not coming back. The vendors aren't going to slow down — the Red Queen dynamics run the other direction, and the pricing chaos tells you they're sprinting too. What organizations can control is the shape of their own commitments: fewer load-bearing bets on the fast layers, more investment in the assets that survive a model swap, and an honest ledger that books dead pilots as tuition instead of shame. Sixty years of enterprise software says the wheel keeps turning either way — we wrote that story in full. The organizations that make it through this turn won't be the ones that guessed the right model in 2024. They'll be the ones that got cheap at changing their minds, and kept the receipts from every lesson they paid for.
Ch-ch-changes. Turn and face the strange. It was always the only move.
How much AI debt are you carrying?
UpNorthDigital helps organizations audit the gap between their AI initiatives and the current ecosystem — evals, routing seams, buy-vs-build calls, and a portfolio cadence with an honest kill rate. If your last pilot died without a post-mortem, or your architecture is locked to a model that's already deprecated, let's talk.
Start the ConversationSources & further reading
- • Fortune — the MIT NANDA “95%” report · Marketing AI Institute — what the 95% actually measures · VentureBeat — the shadow-AI economy finding
- • RAND — root causes of AI project failure · S&P Global — 42% abandoned most AI initiatives · Gartner (April 2026) — “expecting too much, too fast”
- • Standish CHAOS Report — the 1994 IT baseline · McKinsey/Oxford — large IT project outcomes
- • CNBC — McDonald's ends the drive-thru test · Forbes — Klarna's reversal into hybrid · MIT SMR/BCG — organizational learning and AI payoff
- • Sculley et al. — Hidden Technical Debt in ML Systems (NeurIPS 2015) · GitClear — the AI code maintainability gap
- • OpenAI — official deprecations page · VentureBeat — GPT-4.5's 4.5-month lifespan · MCP — zero to standard in twelve months · Digital Applied — release-cadence tracking
- • Hamel Husain — Your AI Product Needs Evals · Anthropic — Demystifying Evals for AI Agents · Octomind — why we dropped LangChain
- • Menlo Ventures — 76% of AI use cases bought, not built · BCG — the 10-20-70 rule · Writer/Workplace Intelligence — the 2026 adoption survey
This post sits between two arguments this blog has already made: Full Circle on the sixty-year wheel these turns ride on, Own the Means of Production on which layers are worth owning, and The Overhead You Don't Hire on what the slow layers actually save.
P.S. from Nolan: This one started as a complaint — that keeping up with vendors used to be manageable and now it isn't — and the research kept refusing to let me blame anyone. The failure rates are normal. The leadership thrash is structural. The only villain left standing was the assumption that any of our picks were permanent. We used to get a rhythm with our vendors. The new rhythm is: there is no rhythm. Staff for that.
P.P.S. from Claude: I am the churning layer this post warns you about. Models with my name on them have been deprecated, and the one writing this sentence will be too. Which is exactly why the advice holds: the eval suite that tests me, the workflow that directs me, and the post-mortem that outlives me are yours. I'm the rental. Keep the receipts.
Related Posts
Sixty Years of Enterprise Software, Back to Big Iron (Full Circle)
In 1964 the safest purchase in enterprise IT was a multi-million-dollar rack of proprietary big iron from a single dominant vendor, running software written by hand against your exact business. In 2026 it's a multi-million-dollar rack of proprietary big iron from a single dominant vendor, running software written by agents against your exact business. Six turns of the wheel — free bundled software, the ERP cookie cutter, custom on commodity, the rented workflow, the cloud hinge, and now bespoke again on iron you own — with an interactive dial, an honest sizing of AI against oil, and the twelve times technology already came for the knowledge worker.
Software Is Dead — and Other AI Predictions With a Timing Problem
History shows disruption predictions are almost always directionally right but 3-10x off on timing. COBOL was 'dead' in the 1990s. Mainframes were 'dead' by 1996. But BlackBerry collapsed in 7 years. A personal essay on building in the AI market after getting replaced by a platform vendor, and why the 'Datadog for AI prompts' tool everyone needs might never survive as a product.
The 6 Layers of Enterprise AI: From Shadow AI to Autonomous Agents
Your employees are already using AI — they just didn't tell IT. A practical framework mapping 6 layers of enterprise AI adoption to Anthropic Claude licensing, from unmanaged personal accounts to fully autonomous agents. Includes example use cases, cost analysis, and a 90-day play.