Big bang rewrite vs incremental replacement of a live system in 2026

TL;DR
A big bang rewrite vs incremental replacement decision in 2026 is a bet on how long your business can wait for a system that does what the old one already does. The rewrite promises a clean codebase on one launch day and usually delivers neither. Incremental replacement, the strangler fig approach Martin Fowler described, swaps one function at a time while the old system keeps running, and moves data from day one.
- IT projects run a mean 1.8 times their estimated cost across 5,392 projects, per Flyvbjerg and co-authors.
- Netscape shipped nothing for nearly three years between version 4.0 and the 6.0 beta while it rewrote from scratch.
- In a 594-comment Hacker News thread on replacing a vendor's system, the most repeated advice was to replace it in pieces and plan the migration first.
- For a system that is live and paying its way: replicate, then migrate, then improve, in that order.
Why the full rewrite keeps getting approved
The pitch always sounds the same. The old system is slow, the vendor is unresponsive, the code is a mess, and a fresh start will fix all three.
The evidence says wait. Flyvbjerg and co-authors analyzed 5,392 IT projects and found the mean ratio of actual to estimated cost was 1.8, distributed as a power law rather than a bell curve. The older McKinsey and Oxford study of more than 5,400 projects found that builds priced above $15 million ran 45% over budget and delivered 56% less value than predicted. A rewrite of a live system is the purest form of that project: large, hard to estimate, worth nothing until the last piece works.
Joel Spolsky called rewriting from scratch the single worst strategic mistake a software company can make, in 2000. Netscape 4.0 shipped nearly three years before the 6.0 beta, and there was no 5.0. His reason still holds: old code has been used and tested, its bugs found and fixed, and it is harder to read code than to write it.
The failure is still being reported first-hand. In November 2025 one engineer described two legacy migrations on Hacker News, both failed: a Java translation nobody had the skills to run in production, and a rewrite that went over budget and was cancelled.
Two forced migrations are already on the 2026 calendar
Google supports the Google Fit APIs only until the end of 2026 and states there is no alternative to the Fit REST API, so any Android app that reads Google Fit is replacing its health data layer this year. Google Play has required apps to target Android 16, API level 36, since August 31, 2026.
Neither deadline asks for a new product. Both ask you to swap one layer inside a system that must keep working, which is incremental replacement in miniature.
How we compared the two approaches
Five criteria, applied to both models.
- Time until new code is in production and paying for itself.
- What breaks if the project pauses at month nine.
- How the data gets across, and who owns that job.
- What happens to feature requests during the replacement.
- Where each model costs more than its plan says.
Published evidence comes from primary sources, linked in the text. Mercury Development sells the kind of engineering work described here, so we scored the incremental model on the same criteria and left its weak spots in.
Big bang rewrite vs incremental replacement at a glance
| Dimension | Big bang rewrite | Incremental replacement |
|---|---|---|
| Best for | Small systems with no users to disrupt | Live systems the business runs on |
| First production value | At launch, or never | Weeks, with the first replaced slice |
| State at month nine if paused | Two systems, one unfinished | One system, partly renewed |
| Data migration | One cut-over weekend | Continuous, from the first slice |
| New features during the work | Tempting, and fatal | Deferred until the replica is live |
| Main risk | Feature parity is never reached | The proxy layer becomes a bottleneck |
| Hidden cost | Running the old system twice as long as planned | Two systems to operate during the transition |
The big bang rewrite
A rewrite is the only model that gives you a clean architecture with no compromises to the old data model. On a small system it is also the fastest route. Microsoft's guidance on the strangler fig pattern says as much: skip incremental replacement when the system is small and replacing the whole thing is simple.
Best for: small systems with a well-understood scope and no live users who would notice an outage.
Facts: value arrives only at cut-over, feature parity is the finish line, and the old system runs at full cost until that day.
The honest minus: a live system's behavior almost never fits on a whiteboard. Fowler's strangler fig article names the two reasons: replacing a serious IT system takes a long time and users cannot wait for new features, and it is hard to figure out the details of existing behavior. Every undocumented edge case becomes a bug against the new build, the business keeps adding requests, and the finish line moves.
Incremental replacement, also known as the strangler fig
Incremental replacement builds the new system around the edges of the old one and retires it piece by piece. Fowler named it after fig vines that germinate in a tree's upper branches, root to the ground and eventually replace the host.
Microsoft describes a facade that intercepts requests to the legacy back end and routes each one to the legacy application or to a new service. AWS frames it as transform, coexist, eliminate. Thoughtworks adds the working steps: define thin slices, introduce an indirection layer, route traffic, retire, iterate.
Best for: systems the business runs on today, where a month without releases costs money and a failed cut-over is unthinkable.
Facts: first replaced slice in production within weeks, old system kept for rollback, investment and returns arrive gradually and visibly.
The honest minus has two parts. AWS warns that the facade can become a single point of failure or a bottleneck, and asks for a rollback plan per refactored service. And you run two systems for the whole transition; Thoughtworks notes the same data may need to stay current in both, which is complex and error-prone. It also has preconditions. If you cannot intercept requests to the back end, or cannot modify the legacy source to redirect internal calls, Microsoft says do not use it. The second condition is where a vendor who never handed over the source becomes a problem.
Build the 1:1 replica before a single new feature
The rule that separates replacements that finish from those that stall: nothing new goes in until the replica does what the old system does.
The temptation is obvious: engineers are in the codebase and the old vendor ignored your backlog for years. Every new feature widens the gap the new system must close before the old one can be switched off. One commenter in the Hacker News thread on replacing a vendor's platform listed the failure modes he has watched: assigning four juniors to replace what an entire company built, and implementing new features during the rewrite. His fix is blunt: never do that, focus on getting a 1:1 replica first.
Parity is also an objective acceptance test. The old system is the specification, and a slice is done when the new path produces the same output for the same input. Run both paths side by side, compare, then route real traffic. Improvement comes after cut-over, and faster than you expect, because by then your team knows the domain better than anyone who has touched the system in a decade.
Data migration starts on day one, not in the last sprint
Code can be rewritten. Data cannot. Every replacement that fails late fails on the data.
The numbers have not moved much in twenty years. Bloor Research reported in 2007 that more than 60% of data migration projects overran on time or budget. McKinsey surveyed nearly 450 CIOs in 2021 and found migrations costing the average company 14% more than planned each year, with 38% of companies delayed by more than a quarter.
On Hacker News, a commenter with a decade in asset management put it plainly: the absolutely critical thing you need to think about from day 1 is the migration onto the new platform. Another, a consultant hired repeatedly to free businesses from vendor platforms, added that the plan has to cover users as well as data, because every replaced piece means retraining.
Mechanically, the first slice you replace includes its data path. The safe way to change a data contract is the parallel change pattern, expand, migrate, contract: add the new structure beside the old, move clients across one by one, then remove the old. Microsoft adds the operational rule: both systems must be able to reach shared data stores for the whole transition. Dual writes, reconciliation reports and a rollback path belong in slice one, not in a hardening phase at the end, and you reconcile one table at a time with the old system still there to check against.
Which approach fits your situation
If you are rescuing a project after a failed vendor or contractor, resist the reflex to start over. You have code that runs, however ugly. Get the source, the repositories and the store credentials into your own name, then replace the worst slice first. If the previous vendor still holds any of those, start with the vendor continuity checklist, and put the right questions to the next team.
If you have outgrown a platform built for a different kind of business, the strangler fig fits directly. Keep the platform as the system of record, put your own layer in front of the parts that hurt most, and move the data out one entity at a time. Get the export terms in writing first, because the platform vendor has every incentive to make that slow.
If the system is small, has no external users and can be down for a weekend, rewrite it.
If the trigger is a platform deadline like the Google Fit shutdown, treat it as a rehearsal for the larger job. Whether the team is in-house or a vendor is a separate question, priced in outsourcing vs in-house development.
What people who lived through a rewrite say
In August 2024 an executive at a logistics company with about 200 employees asked Hacker News whether to bring software development in-house to escape an unresponsive vendor. The thread drew 565 points and 594 comments, and most of the advice was about how to replace the system rather than who should do it.
The pattern is consistent. One engineer who had done this well wrote that going totally greenfield and porting the business over in six months would be misery, that he would build the new system incrementally, and that any monolithic approach to replacing an existing system gives him deep, deep tingly creeps due to past disasters. Another would duplicate one well-understood vertical slice first and abandon the effort if the team cannot replicate even that.
Two more themes recur. The old system is the spec: understand the existing app inside and out before coding the new one. And the failure is political before it is technical: a freelancer said interviewing staff about how each subsystem fits their workflow is what stops the new system copying the old one's design mistakes.
The rewrite you are being sold is usually a handover problem
Opinion, stated plainly. Most rewrite proposals we see are not about the code. They are about the fact that nobody on your side can work on it: the vendor left or never handed over the source, the engineer who understood the system moved on. Starting over is really a proposal to get back control, and a rewrite is the most expensive way to get it. The cheaper way is a team that can read the existing system, run it, and replace it in slices.
We can show that rather than assert it. Element Labs and Barco came to us with configuration software supplied by a hardware manufacturer without source code, so no new features or devices could be supported; we built version 3 on Windows and Mac OS X, and the relationship survived Barco's 2010 purchase of Element Labs. Kensington had an iPhone app built in-house; we extended it to Android in three months and upgraded both in 2015 and 2016. MiMedia first hired us to support legacy Mac and Windows apps, then made us main developer for the next generation. Fitbit had a legacy Windows 8 app that could not sync with wrist trackers; we started in 2011 and still work with the team, now inside Google.
None of those engagements began with a blank repository. Each began with someone else's code, running, with users.
So hold every proposal, ours included, to three questions in writing. What is the first slice, and when is it live. How does the data get across, and how is it reconciled. What does the business have if the project stops at month nine. A phased plan answers each with a date and a mechanism, or it is not a plan.
Written by Rob Devereaux, Chief Operating Officer at Mercury Development. Rob has run the firm's operations from Hudson, Ohio since 2019 and has over 20 years of operational and financial experience. The phased replacement contracts Mercury Development signs, with their slice-by-slice acceptance terms, run through the operations he oversees.
Replacing a system you cannot switch off? Get a phased plan with the data path first
You have the argument. What you probably do not have is a plan that names the first slice, the data path and the rollback. Tell us what the system does, who depends on it and where the source lives today. We come back with a phased plan you can hold us to.