9 min read ·
Price the Business, Then Price the Uncertainty
Use a six-part diligence scorecard to test AI-built software, quantify remediation and delay risk, and choose parity, a discount, or deal terms.

Value a vibe-coded startup from its commercial baseline, then adjust for substantiated technical, legal, security, dependency, and handover risks. Do not apply a standard haircut based on the percentage of AI-generated code: no universal multiple, premium, or discount is established in the available evidence.
Score the six diligence areas; the tool will grade readiness and show which issues a buyer or investor is likely to flag first.
Choose the evidence level that can be verified today. “Planned” work does not count as complete. The score measures diligence readiness, not company value or a fixed valuation discount.
Core evidence is missing. Expect broad technical, ownership, and continuity questions before price confidence.
Architecture Explainability
CriticalTests founder comprehension and whether system knowledge can survive a handover.
- Architecture map and data flows
- Known scaling limits and failure modes
- Recorded tradeoffs and replacement boundaries
Critical-Path Tests
ReliabilityShows whether fast iteration can continue without unpredictable breakage.
- Revenue and access-control workflows
- Integration and data-migration behavior
- Deployment, rollback, and recovery checks
AI-Use and Provenance Log
CriticalConnects generated work to provider terms, human decisions, and review history.
- Tools, accounts, and applicable terms
- Meaningful human edits and approvals
- Prompt or iteration records where relevant
Licenses and Assignments
CriticalUnclear rights can prevent transfer even when the software works.
- Software bill of materials
- Employee and contractor assignments
- Dependency licenses and account control
Security Review
ExposureAutomated scans help, but buyers also need architecture and business-logic review.
- Authentication, authorization, and secrets
- Data handling and dependency exposure
- Incidents, backups, and recovery testing
Bus Factor and Handover
CriticalMeasures whether the company can operate if the original builder becomes unavailable.
- Setup and deployment without the founder
- Defect diagnosis and safe modification
- Rollback and recovery under test
Flags Likely to Surface First
The list prioritizes transfer-blocking rights, key-person dependence, and founder comprehension before ordinary cleanup.
- Licenses and assignments: Complete dependency and license review; resolve contributor assignments and confirm transferability of critical accounts.
- Bus factor and handover: Have an independent engineer set up, deploy, modify, diagnose, and roll back the product without founder intervention.
- Architecture explainability: Document system boundaries, tradeoffs, failure modes, and scale limits; have a nonfounder validate the explanation.
| Score | Readiness Grade | Likely Interpretation |
|---|---|---|
| 10–12 | Ready for focused diligence | Evidence is mostly verified; price discussion can focus on bounded findings. |
| 7–9 | Conditional | Important work remains, but it may be scoped into a reserve or deal term. |
| 4–6 | High uncertainty | Expect a wider range, deeper review, and possible remediation conditions. |
| 0–3 | Not ready | Missing evidence prevents reliable technical and transfer-risk pricing. |
A zero in architecture, AI provenance, licenses, or handover caps the result at Conditional even when the total score is higher.
This is a risk-adjustment framework, not a market-derived valuation standard or a substitute for company-specific financial, technical, security, and legal advice. For a financing, the relevant output may be a pre-money valuation and financing terms. For an acquisition, it may be enterprise value, equity value, cash at closing, or a package of contingent consideration and protections. Those measures are not interchangeable.
Start With the Business, Not the AI Percentage
“Vibe coding” covers materially different practices. Experienced engineers may use AI for implementation while retaining control of architecture, review, testing, and deployment. At the other extreme, agents may generate most of an application with little human inspection. The development method alone does not establish the quality or transferability of the resulting asset.
A largely AI-generated codebase can be modular, documented, secure, and understood by its team. Conventionally written software can be brittle and opaque. Emerging acquisition diligence therefore focuses on whether a company understands, owns, maintains, and can transfer what it built, according to Startups Magazine’s reporting on vibe-coded companies entering sale processes.
Establish the value the company would have if the software had been developed conventionally but had the same customers, economics, market position, team, and technical condition. The baseline depends on stage:
- A pre-revenue company relies on founder capability, customer discovery, evidence of demand, market structure, financing needs, and genuinely comparable financings.
- A recurring-revenue company can be assessed through growth, retention, gross margin, concentration, contract quality, implementation burden, support intensity, cohort behavior, and sales efficiency.
- A mature acquisition target may also support normalized cash-flow analysis and buyer-specific strategic value.
Financings and acquisitions belong in separate comparable sets unless their differences can be normalized credibly. Buyer synergies should also be separated from standalone value.
At the earliest stages, the framework fits early-stage evaluation: product speed matters when it creates validated demand, paid pilots, repeat usage, better positioning, or faster learning. Cheap production of an unwanted product remains cheap production. A technical edge, customer or workflow insight, team capability, and the ability to build something durable remain more informative than code volume or release count, as reflected in Lunera’s investment criteria.
Normalize Costs Before Rewarding Capital Efficiency
Low engineering expense and high reported gross margin may represent genuine operating leverage, temporary founder subsidy, or deferred work. Rebuild the cost structure for reliable operation at expected scale.
The normalized view should include model and API usage, hosting, storage, observability, support, security, compliance, backup and recovery, code review, testing, senior engineering oversight, dependency maintenance, and reliability work. A JPMorgan founder guide similarly cautions that low initial build costs may be followed by inference, observability, security, optimization, training, and commercialization expenses as a company scales in its overview of vibe coding for startups.
| Economic View | What It Shows |
|---|---|
| Reported economics | Historical accounting and current cash use |
| Normalized economics | Sustainable operating, security, support, and improvement costs |
Some costs increase with usage; others arrive in steps when the company needs a larger database tier, security assessment, expanded support, dedicated infrastructure, or senior hire. Classification under an accounting policy should not obscure the cash requirement.
A useful internal measure is sustainable product output per fully loaded engineering dollar. Include review, debugging, refactoring, incidents, support, documentation, infrastructure, and management oversight in the denominator. Measure reliable releases, retained revenue, validated experiments, expansion, and lower support burden in the numerator—not lines of code.
A small team with strong retention, low incident burden, and efficient onboarding may deserve credit. A company that appears lean because one founder works excessive hours and nobody reviews the system has deferred cost, not operating leverage.
Technical Diligence Must Test Comprehension and Transfer
The central technical issue is whether the current core can be maintained and scaled, needs bounded hardening, requires partial replacement, or must be rewritten. A working demo does not establish that the product is secure, comprehensible, transferable, or resilient.
Review architecture and service boundaries, database constraints, critical-path tests, deployment and rollback controls, repository history, observability, authentication, integrations, backups, and known failure modes. Prompt-driven architecture—structure that follows the order in which features were requested rather than a deliberate system design—is a signal to investigate coupling, duplication, and data design, not an automatic rejection.
Founder comprehension carries more weight than the nominal AI share. The team should be able to explain architectural tradeoffs, scaling limits, revenue-critical failure modes, dependency failures, technical debt, incremental replacement options, and the next security investments. An explanation that exists only in one founder’s memory leaves a buyer with key-person risk.
A practical handover test asks an engineer who did not build the product to set it up, deploy it, diagnose a known defect, modify a critical workflow, run the relevant tests, and recover from a failed change. Record elapsed time, missing documentation, founder assistance, and whether the engineer could predict effects on adjacent systems.
Load tests, performance data, database analysis, capacity limits, recovery exercises, and a plan for the next scale threshold provide better evidence than a claim that the cloud will scale. The appropriate standard still depends on the product: customer-facing SaaS needs stronger controls than a disposable proof of concept, while infrastructure and regulated products require deeper professional review.
Replica Tests Locate the Moat Rather Than Erase It
AI has reduced the effort required to reproduce some visible software features. A buyer can use a rough replica to discover what is standard, what requires unusual engineering, and what depends on assets hidden behind the interface.
The Decoder reports that Bain has built rough AI-generated replicas during acquisition diligence to test software capability and the target’s place in the value chain. Its report says a replica influenced at least one bidding decision, but provides no general valuation formula and does not show that replicability alone determines value in its account of replica testing in software acquisitions.
A replica may weaken a claimed source-code moat without reproducing production reliability, proprietary data, integrations, customer trust, workflow embedding, contracts, distribution, compliance, or operating knowledge. The decisive test is whether another team can recreate those assets within a commercially relevant period—not whether it can copy the screens.
Platform dependence also belongs in this analysis. Map which elements the startup controls and which could be impaired by a provider’s price increase, changed terms, native feature, restricted access, degraded model, or discontinued service.
IP, Security, and Dependencies Can Block Transfer
Security diligence should cover authentication, authorization, data handling, secrets management, agent permissions, software composition, known vulnerabilities, incident history, CI/CD controls, monitoring, backup, and remediation ownership. Automated scans identify some exposed secrets, vulnerable packages, and known patterns, but do not replace human review of architecture and business logic.
Human review, software composition analysis, secrets scanning, security testing, software bills of materials, and continuous monitoring are among the controls recommended in Black Duck’s security and governance guidance for vibe coding.
Rights and provenance require separate work. Review code and dependency licenses, employee and contractor assignments, contributor agreements, AI tools and applicable terms, meaningful human edits, architectural decisions, third-party code, and control of critical accounts.
Predominantly AI-generated software may receive limited U.S. copyright protection when meaningful human authorship or creative control is insufficient, depending on the facts and jurisdiction. Records of human edits, review, architecture, and creative decisions can therefore matter, as discussed in Vorys’ qualified analysis of AI-generated software rights.
That does not mean AI-generated code categorically cannot be owned. Copyright, contractual ownership, trade secrets, patents, licenses, account control, and infringement exposure are distinct questions requiring company-specific legal advice.
For each critical model, API, cloud service, database, authentication provider, library, and integration, test outage, price, term, access, and migration scenarios. Determine whether a substitute exists, how long migration could take, and whether the product has a degraded operating mode.
Convert Findings Into Costs, Delays, and Probabilities
Technical debt becomes valuation evidence only after it is translated into company-specific scenarios. Use four cases: maintain as-is, harden, partially refactor, and rewrite. Estimate direct cost, required seniority, timing, probability, operational disruption, management burden, and the effect on sales, retention, expansion, and roadmap delivery.
Expected remediation reserve equals the sum of each scenario’s probability multiplied by its direct cost. Model delay separately through incremental cash-flow or present-value effects. Distinguish permanently lost revenue from deferred revenue, contribution margin from gross revenue, and operating expense from additional financing needs.
Do not subtract a remediation reserve and then impose an unexplained multiple discount for the same problem. A second adjustment is justified only for a distinct effect, such as residual execution uncertainty or permanently weaker margins.
The Worked Example Produces a $3.85M–$5.35M Range
Consider a hypothetical early-revenue SaaS company with $1 million of annual recurring revenue, strong recent growth, good but not exceptional retention, meaningful customer concentration, and a reported gross margin of 84%. Assume assignment-specific acquisition comparables support an illustrative 4.5–6.0 times recurring-revenue range after normalization. This implies a preliminary enterprise value of $4.5 million to $6 million on a debt-free, cash-free basis before transaction expenses.
Diligence then identifies model usage, observability, support, security, compliance, and senior engineering costs that reduce estimated normalized gross margin to 76%. Assume the baseline range already reflects this lower margin, preventing a second deduction.
The technical scenarios are entirely illustrative, not market benchmarks:
| Scenario | Probability | Direct Cost | Delay Effect |
|---|---|---|---|
| Maintain as-is | 20% | $0 | $0 |
| Harden | 45% | $300,000 | $100,000 |
| Partial refactor | 25% | $700,000 | $300,000 |
| Rewrite | 10% | $1,500,000 | $700,000 |
Assume hardening takes three months, partial refactoring six months, and rewriting 12 months. The delay effects represent the valuation-date present value of incremental contribution cash flow lost through postponed sales, expansion delays, additional burn, and management distraction; they are not gross delayed revenue.
The probability-weighted direct reserve is $460,000. The weighted delay effect is $190,000. Subtracting the combined $650,000 adjustment from the $4.5 million to $6 million baseline produces an illustrative enterprise-value range of $3.85 million to $5.35 million.
The arithmetic is not an industry-standard discount. Architecture documentation, critical-path tests, a successful handover, verified assignments, a clean license review, load-test evidence, recovery testing, and provider contingencies could lower the assumed probability of refactoring or rewriting. Unscoped ownership, security, or continuity problems could increase it.
The Verdict May Be Parity, Premium, Discount, or Protection
Parity with conventionally engineered peers is defensible when commercial quality and normalized margins are comparable, the architecture is maintainable, critical workflows are tested, rights are verified, dependencies are resilient, and more than one person can operate the product. In that case, AI-assisted development alone is not a reason for a discount.
A premium is supportable only when the method has produced evidenced advantages: faster customer-validated learning, strong retention and expansion, controlled iteration, sustainable low burn, low incident burden, or durable operating leverage. The premium rewards business performance, not AI usage.
A lower range or remediation reserve is appropriate when the system has opaque architecture, missing critical-path tests, security gaps, uncertain licenses, brittle dependencies, founder-only knowledge, poor reliability, omitted operating costs, or a probable major replacement.
Bounded problems are less damaging than unbounded uncertainty. A known issue with an owner, budget, timing, and remediation plan can be priced. A problem nobody can scope widens the range and may shift risk into closing conditions, escrow, holdbacks, earn-outs, indemnities, retention arrangements, or milestone-based financing.
Rejection-level findings can include an inability to establish rights to a critical asset, unresolved severe security exposure, dependence on inaccessible systems, or no credible way to operate without the founder. The correct output is an explicit range and risk-allocation plan—not an arbitrary “vibe-coding haircut.”