Software Modernization Initiatives That Failed and Why
Most modernization projects fail after launch, not during planning.

Software modernization fails for the same handful of reasons, over and over, across governments, airlines, hospitals, and phone makers. Global IT spending has more than tripled since 2005, from $1.7 trillion to $5.6 trillion in constant 2025 dollars, yet success rates haven't moved much in two decades. Techolution puts the number bluntly: 92% of Fortune 500 companies are running large modernization programs right now, and nearly 79% of them collapse after go-live, not during planning.
That number should stop you for a second. Failure at the planning stage is almost expected, forgivable even. Failure after launch, after the budget's spent and the old system's been switched off, is a different animal. And it's not happening to broke, understaffed teams cutting corners. Many of these are board-mandated, fully funded initiatives with executive sponsorship and plenty of runway. So what's actually going wrong?
The pattern traces back to execution, not technology. The new platform is rarely the problem. How the transition gets scoped, sequenced, and rolled out, that's where things fall apart. Meanwhile, US organizations spend more than $520 billion a year just keeping legacy systems alive, which is exactly why the pressure to modernize never really goes away, even when the last attempt failed. When production-stage collapse happens, it's not an abstract inconvenience: frozen transactions, regulatory exposure, reputational wreckage, and in the worst cases documented below, real human harm.
The technical debt load that makes modernization both urgent and brittle
Every modernization project inherits a pile of technical debt it didn't create. Shortcuts taken a decade ago, undocumented workarounds, and dependencies nobody fully understands anymore have accumulated into the technical debt that the new system inherited, and the new system has to somehow account for it.
McKinsey's research puts technical debt at roughly 40% of the average IT balance sheet. Organizations end up absorbing significant additional costs on top of whatever a modernization project already costs, just to deal with debt that was already there before the project started. Worse, debt isn't just a cost line. Debt doesn't just slow things down. It predicts collapse.
The scale of this globally is hard to grasp. Research across the industry consistently finds that accumulated software debt worldwide represents an enormous remediation burden. The problem has spread well beyond some aging mainframe in a back office somewhere. It's everywhere.
Legacy systems encode decades of business rules, exceptions, and one-off fixes that were never written down anywhere, and none of that appears on a balance sheet. They live in the heads of the people who built the workarounds, and when those systems get replaced, that institutional knowledge often just evaporates. Nobody planned to lose it. It just wasn't anywhere to save.
So why does debt keep piling up instead of getting addressed early? A 2025 survey of more than 500 US IT professionals, run by Saritasa, found that 50% of respondents cited "the current system still works" as the primary reason they hadn't upgraded yet. That's the trap in one sentence. The system works, until the day it doesn't, and by then the debt has compounded quietly for years. Modernization teams never start with a clean slate. They inherit complexity whether they plan for it or not, and any approach that ignores that inheritance just reproduces it in a shinier package.
What the case studies have in common
The cases below span government, healthcare, defense, aviation, and consumer tech. Different sectors, different decades, wildly different budgets, and yet they fail in recognizable, overlapping ways.
These aren't cherry-picked edge cases. They're among the most studied modernization failures on record, which is exactly why they're worth studying again. The threads that keep appearing in these failures: big-bang cutovers with no fallback, dependencies nobody mapped ahead of time, data migration handled as a formality instead of a core milestone, frontline users left out of the design process, no credible way to roll back when things go wrong, and scope that keeps expanding with nobody accountable for it.
Each case below gets its own read. Together, they form something closer to a diagnostic manual.
Birmingham City Council: what happens when you carry old workflows into a new system
Birmingham City Council started replacing its SAP system with Oracle Cloud ERP back in 2018, aiming to modernize finance operations. The system went live in April 2022. By 2025, the cost had climbed to several times its original amount, rising from an original figure to a much larger one, and finance staff were still relying on manual workarounds just to keep basic processes functioning.
The council went on to declare effective bankruptcy. Equal pay liabilities were named as the primary driver of that collapse, with the ERP failure cited as a contributing factor rather than the sole cause.
Here's the root of it: instead of adopting Oracle's standard processes, the council rebuilt its old SAP workflows inside the new system. Over-customization stripped away the very efficiencies the cloud platform was supposed to deliver. Paying for modern software while dragging over the old dysfunction defeats the purpose of switching in the first place; the platform changed, but the actual work didn't get any easier.
Sticking with familiar workflows often gets framed as playing it safe. In practice, it just imports the old system's limitations into the new one, at nearly three times the original price tag.
Canada's Phoenix payroll system: when a go-live date becomes the point of no return
The Government of Canada launched Phoenix on February 24, 2016, meant to consolidate and modernize pay processing for 300,000 federal employees. Within months, systemic errors, underpayments, overpayments, missed payments entirely, hit roughly 80% of the workforce.
The backlog never really cleared. By January 2024, nearly eight years after launch, unresolved transactions still numbered 444,000. The original project was budgeted at C$310 million. By 2018, remediation alone was already running at least C$1.2 billion. As of June 2025, total spending on fixes had passed C$5.1 billion, and compensation payouts to affected employees topped C$711 million on their own.
The core failure here is simple to state and brutal in its consequences: there was no credible rollback plan. Once Phoenix went live with no credible rollback plan, there was no path back to the old system and no way to contain the damage once it started spreading.
Pressure to hit a go-live date and shut down the legacy system is usually organizational, sometimes political, and rarely technical. Phoenix shows what happens when that pressure wins out over actual readiness. Eight years of remediation costs, running at more than sixteen times the original budget, is what gets paid when nobody builds a way back.
Idaho's "Luma" ERP and the NHS: data migration and user adoption as afterthoughts
Idaho rolled out its costly cloud ERP system, known as Luma, in July 2023. Within days, workflows across multiple state agencies collapsed. Payroll and procurement ground to a halt.
The root cause traces back to data. Corrupted records got migrated into the cloud system without proper validation, and the problems only surfaced after go-live, by which point the system was already generating inaccurate payments across the board. Post-rollout, employees struggled to adapt as issues piled up. A subsequent audit found a large sum in undistributed interest, a smaller but still substantial amount in duplicate Health and Welfare payments, and payroll errors spread across the system. Data integrity got treated as a task somewhere on a migration checklist instead of a milestone worth its own scrutiny, and that gap is where the whole thing broke.
The UK's NHS Connecting for Health program tells a similar story at a much larger scale. It launched with an initial budget of £2.3 billion, meant to modernize healthcare IT across a public healthcare system. Poor planning and thin stakeholder involvement pushed costs to £12 billion before the project was largely abandoned. Change management was weak throughout, with no meaningful opportunity to course-correct before the full investment was already committed. Change management was weak throughout, and the clinical staff who'd actually use the systems day to day were never meaningfully brought into design or rollout decisions.
Both cases share the same blind spot. The people who'd use the system, and the data those people relied on, got treated as downstream concerns instead of things worth designing around from day one. Data migration quality and user readiness aren't line items to check off before launch. They're the two clearest signals for whether a system will actually work once it's live.
A large government agency Air Force ECSS and the DoD portfolio: scope expansion without accountability
A military branch's enterprise logistics system, ECSS for short, started in 2004 with the goal of modernizing and transforming logistics operations. It was canceled in 2012 after burning through $1.03 billion over eight years without ever delivering a functional system. The postmortem pointed to flawed program structuring, insufficient upfront planning for migrating data out of legacy systems, and a failure to grapple with integration complexity from the start. Eight years and over a billion dollars, and nothing came out the other end.
Zoom out to a large government agency's broader IT modernization portfolio for fiscal years 2023 through 2025, and the pattern repeats at scale. $10.9 billion got allocated across 24 critical IT business programs. More than half of those 24 ran into serious trouble: inaccurate cost and schedule estimates, scope that expanded without anyone reining it in, weak cloud migration planning, missing cybersecurity strategies, and no required performance metrics to track against. Four major systems alone accounted for 43% of total planned spending, which means a huge concentration of risk sitting inside a small handful of complicated programs. Delays hit logistics and medical support systems directly, so this wasn't abstract underperformance. It disrupted real operations. Plenty of essential systems are still running today without full cybersecurity compliance.
What connects ECSS to the wider DoD portfolio is the absence of any mechanism tying funding to verified progress. Scope kept expanding, timelines kept stretching, and projects kept consuming money without ever having to prove they were moving forward. When oversight gets built around plans instead of demonstrated outcomes, scope creep stays invisible right up until it's too late to fix.
Southwest Airlines and Apple Siri: technical debt deferred until it breaks in public
Southwest Airlines had been running its crew-scheduling system, SkySolver, since 2004, and it was never modernized to handle real peak demand. Winter Storm Elliott exposed exactly how fragile that setup was. The volume of crew reassignments overwhelmed SkySolver completely, and staff had to abandon it for manual processes in the middle of the crisis. More than 16,000 flights ended up stranded. Southwest paid out $600 million in refunds and took on more than $140 million in financial penalties on top of that.
The system worked for nearly two decades. Then a foreseeable stress event hit it, and it didn't.
Apple's Siri tells a quieter version of the same story. Internal reports describe Siri's architecture as fragmented: old rule-based components stitched together awkwardly with newer generative models. That patchwork made updates genuinely hard to ship and produced frequent errors, with internal testing showing core Siri features working correctly only 66 to 80 percent of the time, according to the Software Improvement Group. Between 12 and 15 key AI researchers and executives left Apple within a single year. A "personalized Siri" got previewed in 2024 and never shipped; by March 2025, Apple had publicly admitted the work was taking longer than expected.
Both cases point at the same mechanism. Deferring modernization doesn't make the risk disappear, it just concentrates it in one place, waiting for a storm or a product deadline to expose it at the worst possible time.
The UK Post Office Horizon: when a failing system cannot be replaced
The Horizon software scandal led to wrongful prosecutions of sub-postmasters across the UK, based on accounting errors the system itself produced. Nearly 350 of the accused died, at least 13 by suicide, before receiving any compensation for what they'd been put through.
Two separate attempts to replace Horizon, one with IBM in 2015, another with NBIT in 2024, both failed. The Post Office is still using Horizon today. The government has committed to spending £410 million on a new system, and the Post Office opened bidding for new point-of-sale software in summer 2025, with a decision expected by July 1, 2026.
This case adds something the others don't quite capture. When a harmful system can't be replaced, because every replacement attempt keeps failing, the harm doesn't stop. It just continues. This is primarily a governance and accountability story. It's a governance and accountability story with direct, documented human consequences attached to it. Two failed replacement attempts came before the current procurement effort, and each failure meant more years of reliance on a system whose problems were already well documented. "Modernize eventually" isn't a risk management position. It's a decision, and it comes with a timeline and consequences whether anyone acknowledges that or not.
The four execution gaps that appear in nearly every failure
Strip away the industry and the dollar figures, and four patterns keep recurring across these cases.
Attempting a full system replacement in one go and flipping the switch is the most tempting approach, and often the most dangerous. There's no room to learn and adjust along the way. Edge cases become visible only after launch, when it's already too late to plan around them, and old systems often stay running in the background anyway because nobody's willing to actually turn them off. Incremental approaches, things like strangler-fig patterns or slicing rollout by business domain, consistently beat big-bang timelines on both risk and time-to-value.
Unmapped dependencies and lost institutional knowledge. Legacy codebases carry decades of business rules and exceptions that were never documented anywhere, they exist mostly in the memory of staff who may have already moved on. New teams tend to discover these dependencies mid-migration rather than before it starts, and that's exactly when scope explodes and timelines stretch past anything in the original plan.
Data migration and user adoption treated as afterthoughts. Idaho's Luma system and the NHS program both show what happens when data validation and staff training get pushed to the end of the checklist instead of built in from the start. By the time problems become visible in the rollout, the system's already live and already producing bad output.
Scope expansion with no accountability tied to funding. The DoD portfolio and ECSS both show what happens when oversight tracks plans instead of verified progress. Money keeps flowing, scope keeps growing, and nobody notices the drift until the whole thing is unrecoverable.
None of these gaps are about the technology chosen. Oracle, SAP, custom-built government platforms, it doesn't much matter. The failures above happened in the sequencing, the rollout, and the decisions about who got a say before launch day. That deserves the longest attention.


