AI-lens technical due diligence
Diligence for the AI era: a method, not a verdict.
We publish the method because the discipline is the point – you should be able to judge its rigour without ever seeing a client's data room.
Surface and substrate
AI-assisted development has made large parts of most software products replicable at a fraction of their original cost. A diligence method needs a way to separate what AI has made cheap to copy from what it has not touched at all.
Surface
The interfaces, dashboards, workflow forms, onboarding journeys and reporting views a user sees. Modern AI-assisted engineering reproduces most of this quickly.
Substrate
Regulated integrations, live counterparty networks, accumulated data and operating history, embedded adoption, and the switching costs all of that creates. If a moat holds, it holds here.
“Hard to build” is not the same as “exclusive.”
Read the full method →Every finding carries a confidence flag
Evidenced
Directly supported by data-pack evidence sighted.
Partially evidenced
Some but not all supporting evidence sighted, or strong, specific public-domain corroboration short of conclusive.
Inferred
A public-domain hypothesis or sector pattern with no direct corroboration.
The flag is not a hedging device appended to soften a finding. It is the finding.
Read the full method →What we refuse to do
A method is judged as much by what it declines to assert as by what it concludes. The four refusals below are where a less rigorous process quietly fills the gap with confidence the evidence doesn't support.
No promised return on investment
Every value figure is an indicative range, stated with its assumptions, offered for validation – never a warranted saving.
No exclusivity verdict without the evidence it requires
Where the contract or management evidence is not available, the finding is stated as a hypothesis, the gap is named, and the shortest realistic path to closing it is set out.
No smoothing over a gap
Where the evidence stops, an itemised evidence-gap register says so.
1. The problem with most AI diligence
Two failure modes dominate technology diligence in the AI era, and they sit at opposite extremes. The first is the checklist inherited from the last cycle, applied as if AI-assisted development were a marginal efficiency gain rather than a structural change to build economics. The second is the opportunity list dressed up as a value case – plausible AI applications any experienced technologist could generate from public materials alone, presented with a confident number attached.
A serious method holds every claim against the honest base rate for AI value creation:
MIT's NANDA initiative reports that '95% of organizations are getting zero return' from generative-AI investment, and that 'just 5% of integrated AI pilots are extracting millions in value, while the vast majority remain stuck with no measurable P&L impact.'
Partially evidencedThe linked copy is an unofficially-mirrored one — no publisher-hosted URL could be located — of a report the authors themselves label "preliminary findings"; the flag also reflects a disclosed but preliminary, convenience-sample methodology (52 interviews, 153 conference-survey responses, 300 public deployments reviewed).
MIT NANDA (Project NANDA, MIT Media Lab), The GenAI Divide: State of AI in Business 2025 (2025)
RAND's 2024 report states that 'by some estimates, more than 80 percent of AI projects fail ... twice the already-high rate of failure in corporate information technology (IT) projects that do not involve AI' — a figure footnoted to a Fortune commentary piece, not to RAND's own data; RAND's own contribution is 65 structured interviews into why AI projects fail, not a measurement of this figure.
InferredA widely repeated public-domain estimate with no primary methodology we could sight behind it — the report itself footnotes the figure to a Fortune commentary piece, not to RAND's own data.
Gartner forecast in July 2024 that 'at least 30% of generative AI (GenAI) projects will be abandoned after proof of concept by the end of 2025' — a prediction whose horizon has now passed, with no published outcome confirming or revising it.
InferredA forecast is a hypothesis about the future by definition, and this one remains unconfirmed either way: its horizon has now passed with no published outcome confirming or revising it.
McKinsey's 2025 global survey of 1,993 respondents across 105 countries found that '88 percent report regular AI use in at least one business function,' while 'about 6 percent' qualify as 'AI high performers' — respondents attributing 5 per cent or more of EBIT to AI use who also report 'significant' organisation-wide value from it.
EvidencedChecked against the publisher's own report, with a disclosed sample and a stated definition behind the headline figure.
McKinsey & Company (QuantumBlack), The State of AI: Global Survey 2025 (2025)
None of these say AI is unimportant to an investment case. They say the gap between a plausible AI story and a proven one is where most diligence quietly fails.
2. Surface versus substrate
The core analytical frame we use starts from a simple observation: AI-assisted development has made large parts of most software products replicable at a fraction of their original cost and time. That changes what “defensible” means, and a diligence method needs a way to separate the part of a business that AI has made cheaper to copy from the part it has not touched at all.
We call the first the surface: the interfaces, dashboards, workflow forms, onboarding journeys, reporting views and document templates a user actually sees. Modern AI-assisted engineering can reproduce most of this quickly, and increasingly cheaply, for almost any category.
We call the second the substrate: the things underneath the interface that a well-funded competitor cannot simply generate – regulated integrations into third-party systems, a live multi-party network of counterparties, accumulated data and operating history, embedded customer adoption, and the switching costs all of that creates. If a moat holds, it holds here.
The diligence value is not in describing the substrate – most competent technologists can list “integrations, data, network effects” for any asset. It is in testing whether the substrate is genuinely defensible or merely assumed to be. “Hard to build” is not the same as “exclusive.” An integration that took years to build is still substitutable if a competitor can obtain equivalent access on a comparable timeline; a network is still a moat only if switching away from it is genuinely costly for the counterparties inside it. The method has to establish, rail by rail, which of those is actually true – and how long, and with what access, a well-resourced competitor would need to reproduce equivalent reach. It also has to name the specific, realistic actors who could plausibly try – resourced incumbents, adjacent players moving in-house, AI-native clean-slate entrants – and distinguish who could build a lookalike product from who could actually threaten the position. Those are very different lists, and conflating them is one of the more common ways a diligence view overstates risk.
3. The confidence flag – the deliverable, not the decoration
Because that test rests on contract and relationship evidence as much as on technical evidence, and because most value-creation estimates rest on operational data a diligence team may or may not be given, every finding we produce carries one of three confidence flags:
- Evidenced – directly supported by data-pack evidence sighted: documents or data actually reviewed.
- Partially evidenced – some but not all of the supporting evidence sighted, or strong, specific public-domain corroboration that falls short of conclusive.
- Inferred – a public-domain hypothesis or sector pattern with no direct corroboration, held pending clarification.
This is not a hedging device appended to soften a finding. It is the finding. A checklist that says “the integration is defensible” and a report that says “the integration is defensible – Evidenced, against the partner agreement sighted” are making different claims, and only one of them can be relied on in an investment paper. The flag tells the reader, at a glance, what is established fact, what is directionally supported, and what is a reasoned hypothesis still awaiting evidence – for every rail, every opportunity, every value range, not just the headline conclusion.
The same discipline applies on the upside side of the lens. An AI value-creation opportunity is easy to generate; almost any document-heavy, workflow-heavy business will yield a plausible list. What is hard, and what the flag exists to make visible, is which of those opportunities the underlying data actually supports – data quality, integration maturity, workflow standardisation, real volumes and a real cost base – versus which rest on a sector pattern that sounds right. A high-value opportunity sitting on evidence that is merely Inferred is scored and reported as exactly that, not smoothed into a number that looks more certain than it is.
4. The confidence ladder
Diligence is rarely a single pass. Questions get deeper as access improves – a data pack expands, management becomes available, a document request is answered. The method has to describe that progression honestly, and the honest description is a ladder, not an upsell.
Each stage of deeper work is scoped to a named question, and framed by what it is expected to move – which specific findings it should shift from Partially evidenced or Inferred to Evidenced, and why. It is never sold as generic “more accuracy for more time.” If a question can only be closed by a specific piece of evidence – a contract, a management session, a volume dataset – the method says so, and names exactly what closing it would take.
The corollary is an evidence ceiling, stated plainly rather than discovered awkwardly later. If the evidence a finding needs is not available – no contract in the pack, no management access granted – further analyst time on the same materials cannot manufacture that evidence. A rigorous method says when it has reached that ceiling and recommends stopping, rather than continuing to extend a fee against a gap that more hours cannot close. That is a harder thing to say than “let's do more analysis,” but it is the difference between a method that serves the investment decision and one that serves its own billings.
5. What we refuse to do
The discipline shows most clearly in what the method will not produce, because these are the exact places a less rigorous process quietly fills the gap with confidence the evidence doesn't support.
We never state a promised return on investment. Every value figure is an indicative range, stated with its assumptions, offered for validation – never a warranted saving or a point figure an investment committee can bank on without independent confirmation.
We never issue a firm exclusivity verdict – “this relationship cannot be replicated” – without the contract or management evidence that verdict actually requires. Where that evidence is not available, the finding is stated as an evidenced hypothesis, the gap is named, and the shortest realistic path to closing it is set out. A confident answer to a question the evidence cannot yet support is not rigour; it is the opposite of it.
And we never smooth over a gap. Where the evidence stops, the method says so, in an explicit, itemised evidence-gap register naming what could not be established and what would close it – rather than folding the uncertainty into upbeat prose and hoping the reader doesn't notice the load-bearing assumption underneath.
6. Why this beats a checklist
A checklist is a good instrument for confirming that known boxes are ticked. It is a poor instrument for an investment question that turns on judgement under partial evidence – which is what an AI-era moat and value case both are.
A checklist cannot tell you which integration rail is genuinely exclusive and which is merely long-standing; both look identical on a checklist that only asks “is there an integration.” It cannot tell you which AI value-creation opportunity will evaporate on weak data foundations, because a checklist scores the opportunity, not the readiness underneath it – and the base rate in §1 exists precisely because readiness, not opportunity, is what usually kills the value case. And a checklist has no native way to say “we don't know yet, and here is exactly why” – it either ticks the box or it doesn't, with no register for the honest middle ground where most real findings actually sit.
The confidence flag, the surface/substrate frame and the evidence-gap register exist to hold that middle ground explicitly, so a reader can tell a genuinely tested finding from an assumed one at every level of the report – not just in the executive summary.
7. Who this is for
This method is built for people who have to defend a number in a room they don't control: PE deal teams building the investment case, and investment committees deciding whether to back it. Both need to know, precisely, which parts of an AI-era thesis are established and which are still assumed – not because uncertainty is a failure, but because knowing exactly where it sits is what makes a decision defensible. A method that tells you that, honestly, at every finding, is worth more than one that tells you everything with equal, unearned confidence.
Koralis AI applies this method as part of technical due diligence, working alongside lead advisers on AI-era moat and value-creation questions. Findings referencing sensitive personal or financial data are assessed with GDPR-aware handling in mind; this is a diligence discipline, not a claim of legal compliance. This note describes method only; it does not reference, and should not be read as referencing, any specific engagement or client.
GDPR-aware handling; not legal compliance advice.
Assessing a technology asset?
We work alongside lead advisers on AI-era moat and value-creation questions.