Best Firms to Take a Vibe-Coded App to Production (2026)
September 30, 2026 by Paul Byrne
The honest rewrite vs refactor answer is that you refactor by default. A full rewrite is correct in two situations only: the data model cannot express what the business sells, or the architecture has a ceiling that gets worse as volume grows. Everything else, including code you find unreadable, is a refactoring problem you can schedule.
That holds whether it came from an offshore team, a departed developer, or an AI tool.
Score your own system on the table below. If you land in the middle, book a free 30 minute scoping conversation and we will read the codebase with you.
Mark every row that describes your system today. They are ordered by how strongly they bind: one from the top outweighs five from the bottom.
| Signal you can observe | What it actually means | Points to |
|---|---|---|
| A new feature needs a column meaning different things per customer | Data model cannot express the business | Full rewrite |
| Built for one organization, now must serve several, data isolated | Data model cannot express the business | Full rewrite |
| Load doubles, response time more than doubles | Architectural ceiling, worsens with volume | Partial rewrite |
| A nightly import that finished in an hour no longer finishes | Ingestion ceiling in one subsystem | Partial rewrite |
| One subsystem causes most incidents, the rest is quiet | Localized failure, not systemic | Partial rewrite |
| Every change breaks something unrelated | Coupling plus missing tests | Refactor, tests first |
| Inconsistent naming, long files, mixed style | Readability: costs time, not capability | Refactor |
| Nobody can explain what a module does | Knowledge gap, not a verdict | Neither yet, document first |
| Stable, no complaints, nothing revenue bearing blocked | Uncomfortable, not constrained | Do nothing yet |
The two options fail in different ways, which matters more than the headline price.
| Refactor | Rewrite | |
|---|---|---|
| Cost profile | Incremental and stoppable. Spend a quarter, stop, keep the gains. | Committed up front. Worth nothing until it reaches parity. |
| Time to first value | Days to weeks per improvement. | Months before a user sees anything. |
| Main risk | You spend, and still hit the same ceiling later. | You rebuild every undocumented rule, and find the missed ones in production. |
| Existing users | Keep their system, see it improve gradually. | Wait, then migrate. Migration is its own project. |
| What you end up owning | The same system, cheaper to change, original constraint intact if structural. | A system shaped for today, plus an old one to retire. |
A ceiling gets worse as volume grows and a bug does not. Test it without reading code: take a slow operation, compare it to how long it took at half your current volume, and project double. A flat curve means work to do but no ceiling. If cost or latency climbs faster than the volume causing it, adding customers makes it worse. The clearest case is the job that used to finish and now does not, like an export that went from minutes to hours to overnight. Ordinary bugs are constant, at the same rate with ten users or ten thousand.
No, and it is the most expensive wrong reason in this decision. Joel Spolsky made the case in Things You Should Never Do, Part I in 2000: code looks worse than it is because reading code is harder than writing it, and the ugliness is often accumulated handling of real cases that a rewrite rediscovers one production incident at a time.
Also note who is talking: whoever calls the code unreadable is usually whoever would be paid to replace it. “I would not have built it this way” and “this cannot do what the business needs” are different statements, and only the second is a business case.
More than the build, because the build is the part you can estimate. The real cost is the feature freeze: while the new system chases parity, the old one stops improving, so your roadmap stalls for the full duration. Add migration and a period running two systems, and a rewrite priced as a build routinely lands at two to three times the quote.
The worst outcome is the abandoned rewrite: eighteen months in, the budget is gone, parity is not reached, and you run the original plus a codebase nobody will finish. A refactor that turns out wrong leaves a cleaner system you still run. A rewrite that turns out wrong leaves nothing you can deploy.
It is the middle path: put a router in front of the old system and replace one piece at a time behind it, so new code goes live in weeks, not quarters. Martin Fowler named the pattern in Strangler Fig, and Microsoft documents the mechanics and limits in the Azure Architecture Center’s Strangler Fig entry. Both make the same point: incremental replacement lowers migration risk and returns value before the migration ends.
Buyers skip it because it is harder to explain than “we are rebuilding it,” and because it costs more in discipline: two systems side by side, data consistent across both, no declaring victory at 80 percent. Microsoft is clear it does not always fit: it needs you to intercept requests and change the legacy source.
You assess the system independently of whoever built it, because that is the normal case, not the hard one. What replaces the missing person is evidence: the schema, the traffic, the error logs, the deploy history, and behavior under real load. First, write characterization tests. They record what the system currently does, including the parts that look wrong, so any later change tells you what it broke, and a rewrite needs that record to prove parity.
Razoyo’s codebase rescue work runs seven production readiness checks on every takeover: security and access, data integrity, architecture, test coverage, observability and alerting, a real deploy path, and documented ownership. A fixed list covers what was not visibly broken, which is where inherited systems fail.
The evaluation is the same, but two weaknesses show up more often: thin test coverage, and a data model designed for the demo rather than the business. The tool optimized for something that runs, not something that holds, and those are the two dimensions this decision turns on.
That is not a verdict against the code. A prototype that proved the idea did its job, and much of it is reusable. It means hardening an AI-built app for production starts with the schema and the tests, not the features. Whether it came from an AI tool, a contractor or a stalled agency, the assessment does not change.
Buy the recommendation separately from the build, and own the document. If one engagement produces both the verdict and the work, the verdict has a commercial interest in one answer.
That is what Razoyo’s System Spec is for: $2,500 fixed, covering goals, constraints, core flows and success criteria, and producing a documented build or do-not-build recommendation with a senior engineer estimate. It is yours whether or not you continue, so you can take it to any team for a competing bid. It is step 1 of Razoyo’s published process: free consultation at $0, then Requirements Package at $2,500 fixed, a working MVP from $5,000, MVP to production from $10,000, and ongoing support from $1,500 a month. On build work, once fees are paid you own the code we write for you.
The test of an honest recommendation is whether the firm will tell you to do nothing.
Five questions, each scored 0, 1 or 2. It is the structure used inside a System Spec.
1. Data model. Can the schema express what the business sells today without workaround columns or duplicate records? 0 = yes. 1 = workarounds, but contained. 2 = the business has outgrown the schema.
2. Ceiling curve. When volume doubles, what happens to cost and latency? 0 = double or less. 1 = degrades, with headroom. 2 = more than double, hurting now.
3. Change cost. How long does one small change take to reach production, and how often does it break something else? 0 = days, rarely. 1 = weeks, sometimes. 2 = nobody will touch it.
4. Knowability. Can anyone say what the system does without reading all of it? 0 = yes. 1 = partly. 2 = no.
5. Business pull. Is there something revenue bearing you cannot do until this changes, with a date attached? 0 = no, this is discomfort. 1 = something is slowed. 2 = a named contract or launch is blocked.
Bands
Two override rules keep this honest. If question 5 scores 0, the answer is “do nothing yet” whatever the total: an 8 from questions 3 and 4 alone is unpleasant, not failing. And a full rewrite needs a 2 on question 1 or 2, because ugly code, missing tests and absent docs can reach 6 points without any being a reason to start over.
Matrix is the clearest case in Razoyo’s own work of a system that genuinely needed the rewrite. It began as an internal hotel sales CRM built by an outside firm, with poor performance and an interface that was not sellable. On question 1 it scored a 2: the original was built for a single organization, so every piece of state, every data model and every workflow was scoped to one tenant. The plan was to sell it to other hotel groups, and no refactoring makes a single-tenant schema multi-tenant.
Razoyo rebuilt it from the ground up as a multi-tenant SaaS on Elixir and Phoenix LiveView, with data isolation, role based access and the infrastructure to onboard new customers. It launched as M1 Intel and now runs hundreds of properties, with Razoyo as the team that runs it. See the full Matrix case study. The call was correct because questions 1 and 5 both scored a 2, which is rare.
Spider, the platform behind Lighthouse’s hotel business intelligence, is the other shape: an earlier internal tool built by an outside team that hit a hard scaling ceiling, with reservation exports too large and daily ingestion failing. That is a question 2 score of 2. Razoyo rebuilt it as Spider, the hotel business intelligence platform behind Lighthouse, which placed #1 of 74 for Business Intelligence at the 2025 HotelTechAwards.
If you are in the 3 to 8 band, the next step is not a build quote but an assessment you own, from someone not bidding on the outcome. Razoyo has taken inherited systems into production since 2011, from The Colony, Texas, for clients across the United States, through custom systems and software rescue work, and on our own product AutomaticFFL.
Book a scoping conversation and we will tell you which band you are in, including when the answer is that you should not spend the money yet.
Refactor. Changing how code is written without changing what it does, so it is cheaper and safer to modify.
Full rewrite. Building a replacement from scratch, migrating users and data, and decommissioning the original.
Partial rewrite. Replacing one component while the rest of the system keeps running.
Strangler fig. Incremental replacement: a router in front of the old system sends a growing share of requests to new code until the old one is switched off.
Data model. The structure deciding what facts your system can record at all. Changing it is the most expensive change there is, which is why it drives this decision.
Multi-tenant. Several organizations share one system with their data isolated. A single-tenant system assumes exactly one organization.
Scaling ceiling. A limit where cost, time or failure rate grows faster than the volume causing it, so the system worsens as the business grows.
Characterization test. A test recording what the system currently does rather than what it should do, so any later change shows what it altered.
Feature parity. The point where a replacement does everything the original did, which rewrites miss because the original keeps adding features.
Should I rewrite or refactor an inherited codebase? Refactor by default. Rewrite only when the data model cannot express what the business sells, or the architecture has a ceiling that worsens as volume grows. Ugly code and missing tests are refactoring problems, and if only one subsystem has a ceiling, replace that subsystem alone.
How do I know if my system has a scaling ceiling? Compare the same operation at half your current volume and project it at double. If cost or latency grows faster than the volume causing it, that is an architectural ceiling. If the operation is slow at every volume, that is a bug, which is far cheaper to fix.
Is it cheaper to rewrite or refactor? Refactoring is almost always cheaper, and it is stoppable: spend for a quarter, stop, and keep what you gained. A rewrite is committed up front, is worth nothing until it reaches parity, and its real cost includes the feature freeze, the migration and running two systems at once.
What if the original developer is gone and left no documentation? That changes the sequence, not the decision. Start with characterization tests that record what the system currently does, including the parts that look wrong, so any later change tells you what it broke. A rewrite needs that same record to prove it reached parity.
Does a rewrite make sense for an AI-generated prototype? Sometimes, but for the same reasons as any other system, not because of the tool that built it. AI-built prototypes tend to have thin test coverage and a data model designed for the demo, so they fail the data model test more often than average.
How do I get an independent recommendation instead of a sales pitch? Buy the assessment separately from the build, and make sure you own the document. Razoyo’s System Spec is $2,500 fixed and produces a documented build or do-not-build recommendation with a senior engineer estimate, yours whether or not you continue. An honest assessment will sometimes tell you to do nothing yet.
These links open AI platforms with pre-written prompts about this page.
We use cookies to improve your experience. Do you accept?
To find out more about the types of cookies, as well as who sends them on our website, please visit our cookie policy and privacy policy.