How to Evaluate an AI Product Before an Acquisition
Learn how to evaluate an AI product before an acquisition by testing technical ownership, inference economics, data rights, regulatory exposure, and post-close operability.
Date
August 31, 2026
category
Artificial Intelligence
READ
8 min read

One test predicts most of what a demo hides: could your own engineers run, retrain, and defend the system the day after close? Every finding here traces back to gross margin, defensibility, and the price you pay. By the end, you'll turn a demo into a go, reprice, or walk decision.
The one question that reframes the whole evaluation
When you evaluate an AI product before an acquisition, one question predicts most post-close surprises. The day after close, can your team run and retrain the system, and still defend it to a customer or a regulator? If the answer depends on the founder's laptop or a contract that expires at closing, you've found risk a demo will never surface.
Most due diligence splits into silos. Lawyers review contracts and engineers review the codebase, while the deal team builds the valuation model on its own. The three reports rarely reference each other.
That gap is where the real problems hide, because a technical fact only matters once you can price it. So hold every finding to a chain that runs from operational effect to business impact to a deal action. A missing eval harness means you can't verify accuracy before shipping, so you reserve budget to rebuild it and lower your offer.
"AI-washing," where a company overstates how much real machine learning sits under a product, is the symptom people fixate on. The diagnosis is narrower and more useful: can you own this system, and what does owning it cost? Keep that question in front of every session, and the rest of the evaluation organizes itself.
Wrapper or real capability: what you're actually buying
An AI product stacks three layers: a data pipeline that feeds it, model development that shapes it, and serving plus monitoring that keep it running. Ask which layers the company built and which it rents. A thin product often turns out to be a prompt and a web form on someone else's API.
Request specific artifacts. Ask for a data-flow diagram, model cards for anything they claim to have trained, an eval harness or dashboard, and a year of incident logs. Each one shows whether the capability exists as engineering or as a story.
Some tells are hard to hide. With no eval harness, no one can prove the model's quality held up after an update.
With no fine-tuning or retrieval layer, the "proprietary model" is really a system prompt. With no incident logs, either nothing runs in production or no one is watching.
A good answer sounds specific: "here's our eval set, here's accuracy by customer segment, here's the drift alert that fired in March." An evasive answer stays abstract, or points to the founder as the only person who understands the pipeline. Specificity under pressure is the signal you want.
Treat foundation-model dependency as a spectrum. Building on a hosted API is fine and often smart, though it changes what you're buying and how defensible it is.
A company that fine-tunes open weights and owns its retrieval layer sits far from one that just forwards requests to a single vendor. Price the position, don't just note it.
Inference economics: the margin question hiding in the demo
The demo won't show you the bill. To find the real margin, pull twelve months of cloud and model-provider invoices and separate inference spend from the rest. Inference is cost of goods sold rather than research, so it belongs in COGS once the product serves paying customers.
Rebuild the unit economics from those invoices: cost per customer per month, then how that cost moves as usage grows and prices change. Providers cut prices often, but they also deprecate models and raise rate limits, so test both directions. A product that only works at today's token price has a fragile margin.
This matters because margin sets the multiple. A durable 70% gross margin supports a very different revenue multiple than a 40% margin that erodes as usage climbs. When you reprice inference honestly, you reprice the whole deal.
Watch for token and GPU cost booked as research and development instead of COGS. That single reclassification can turn an apparent 80% gross margin into something far lower once the product is yours.
The distance between the pitch and the operated reality is the gap between what passes evaluation and what survives production. If the target needs stronger production architecture, AI automation and production engineering can help turn prototype behavior into an operable system. So ask for the invoices, not the summary, and rebuild the number yourself.
Can the model actually be run and retrained after close?
Running the system tomorrow depends on rights you might not be buying. Many AI companies train on customer data under master service agreements, and some of those licenses end on a change of control. Read the survival and assignment clauses before you value the training data, because the rights can die at closing even though the data stays put.
Then separate benchmark accuracy from production accuracy. Vendors quote the benchmark, while customers live with production, where messy inputs pull quality down.
As an illustrative example, not a real client, a model advertised at 98% benchmark accuracy can land near 71% on live traffic. Without drift monitoring, no one notices until a customer does.
Ask how the team rebuilds a model from source artifacts. Mature MLOps means you can reproduce a deployed model from versioned data and code without the person who first trained it. If reproduction lives only in one engineer's memory, you're buying a dependency, and you should tie retention to earn-outs.
In one MIT Sloan study, 33% of acquired employees left within the first year, compared with 12% of similar regular hires, so plan for that rather than hoping.
Finally, check where personal data physically lives. GDPR's right to erasure (Article 17) lets someone demand deletion of their data. When that data is baked into trained model weights, honoring the request can mean retraining the model rather than deleting a row. That work belongs in your cost model before you sign.
Each of these findings ends the same way. It becomes a number in the valuation, a clause in the contract, or a line in the integration budget.
Regulatory and enforcement exposure you inherit at close
Whatever the target has shipped, you inherit its regulatory exposure at close. A few exposures deserve a hard look, because each one can force spending or destroy an asset you thought you bought.
Start with the EU AI Act, which entered into force on 1 August 2024. Obligations for general-purpose AI (GPAI) models became applicable on 2 August 2025, and the Commission's enforcement powers over GPAI providers apply from 2 August 2026.
Under Article 101 of the AI Act, penalties for general-purpose AI providers can reach up to 3% of global annual turnover or €15 million, whichever is higher. The law sorts systems into risk tiers (prohibited, high-risk, limited, and minimal). Confirm where the product sits under the EU AI Act's risk tiers and general-purpose AI obligations.
Open-weight models carry their own lock-in. Under Meta's Llama license, a licensee whose products exceed 700 million monthly active users in the preceding calendar month must request a license from Meta. Meta may grant or refuse that request at its sole discretion.
If the target builds on Llama, your own user base can push the combined entity past that threshold after close. The license you relied on then stops being automatic.
Regulators can also order destruction of the model itself, not only the data. In the FTC's Rite Aid order from 2023, the company had to delete unlawfully used data and destroy any models or algorithms derived from it.
Rite Aid was also banned from facial recognition for five years. An algorithm you're paying for today can become a legal liability tomorrow, so ask how the training data was collected and consented.
Add the US state-law patchwork on top. States are passing their own AI and privacy rules at different speeds. A product compliant in one state may not be in another, and that fragmentation raises your compliance cost.
Map the target's markets against the rules that apply there before you model integration.
From findings to a go / reprice / walk decision
Now join the lenses. Score each finding on severity across the technical, commercial, and legal reads, then sort it into a bucket. Small, fixable issues become a reprice.
Serious but bounded risks you protect against with reps and warranties, an indemnity, or an escrow. A finding you can't price or protect against is a reason to walk.
Reserve post-close capital before you sign, not after. Budget for retraining models whose rights or accuracy are shaky, for retaining the engineers who can reproduce them, and for compliance work regulators may demand. Deals rarely fail because a risk existed; they fail because no one funded the fix.
The build-versus-partner call sits right here. If the target's system needs a rebuild, the honest question is whether your team can actually own and operate the codebase after the founders leave. Custom software development can support that rebuild when the acquired product needs architecture, integration, or operational hardening after close. Answer that before you set the earn-out, because it changes both price and integration.
Keep the discipline simple: every finding lands in the valuation model, the reps, or the integration plan. Nothing stays as a note.
If you want a second set of hands, Scaylar offers a focused technical read of an AI target's architecture and inference economics before you sign. You go into the negotiation with the numbers already rebuilt.
Conclusion
The test never changes: can your engineers run, retrain, and defend the system the day after close? The acquirer who prices that reality, with clear eyes on what it costs to own, pays the right number, while the one who trusts the demo overpays for a story. The difference comes down to discipline. Run every finding down to a number in the valuation model or a specific deal action, and the price you offer will reflect what you can actually own and operate, not what the pitch promised.


