This is an abridged edition of my paper published by the Center for Trustworthy AI. The full edition, including the evidence stack, the division of responsibility between technologists and counsel, the appendix, and a downloadable PDF is at centerfortrustworthyai.org/aidiligence.
Since information technology time immemorial—let’s call it fifty years, give or take—the IT group has played, at best, a supporting role in the valuation and viability of a typical organization. When a merger, acquisition, investment, or other financial transaction was at play, a buyer’s questions about their target’s technology arrived through familiar channels. Intellectual property. Privacy. Cybersecurity. Contracts. Whatever the technology was, it could be routed into one of those lanes and examined by someone who had examined a hundred like it.
Artificial intelligence does not route cleanly, and AI itself has irrevocably broadened the field of organizations impacted by these sorts of legal diligence questions far beyond private sector firms engaged in a financial transaction.
AI and the data estate that underpins it have become core strategic drivers for nearly every organization, both in the private and public sectors. Further, AI is inherently non-deterministic. In traditional software, a passing test result indicated that the system performed as expected. In AI, a passing test result indicates that the AI system performed as expected on that instance of the test. It does not guarantee future successful runs. In a world of non-deterministic technology, past performance is not a guarantee of future results. Finally, a model can drift without a single line of code changing. Training data carries obligations the organization may never have inventoried. An agent built by an enthusiastic colleague may be running in production without anyone in IT having been aware of it.
For these reasons, nearly every organization now finds itself in one or more of three scenarios.
Scenario One — In the tunnel, wherein diligence ahead of a major financial transaction is weeks away or underway. This is the typical mode in which investors, counsel, and organizations engaged in such a transaction are operating.
Scenario Two — On the horizon, wherein a transaction is plausible within six to eighteen months. These organizations have breathing space. Barely.
Scenario Three — Steady state, the baseline condition of every organization, including those in the first two. No transaction need be anticipated, or even possible, as is the case in public sector organizations that are neither bought nor sold.
Admittedly, scenarios one and two apply to a fraction of the world’s institutions. Scenario three applies to them all.
The legal profession has noticed, and it has responded the way it responds to any durable change in risk, by writing down what to ask.
The clearest statement of that checklist comes from Danny Tobey, Sean Fulton, and Coran Darling of DLA Piper, whose December 2025 framework for AI-specific diligence sets out, area by area, what buyers and their counsel now look for in a target’s AI. Proprietary development. Third-party systems. Deployment and controls. Training data and its governance. Generative AI and its outputs. The regulatory landscape wrapped around all of it, and the organizational governance meant to hold it together. The framework is thorough, it is in use, and it is not going away.
Danny Tobey, Sean Fulton, and Coran Darling, “Digital Diligence: AI-Specific Diligence and Subject Matter Experts in Corporate Transactions,” DLA Piper with Practical Law Intellectual Property & Technology, 8 December 2025.
Which means the question for technology and organizational leaders alike has changed. It is no longer what will they ask. It is can you answer, whether in the case of due diligence, routine audit, or in the aftermath of a calamity.
Put plainly, AI diligence is an assessment of AI maturity performed by strangers, under time pressure, with your valuation, legal status, or existence at stake. The strangers are competent, they are motivated, and they arrive with a list. What they find sets the terms of whatever comes next. In a transaction, that is the price, the structure, how risk is allocated between the parties, and how much the buyer holds back to fix after close what it found before it. In an audit or an investigation, it is the scope of the inquiry, the remedy, and whether the matter ends there. In the aftermath of a failure, it is whether the organization can show that it acted reasonably, which is usually the only question that matters. Those are the stakes the legal framework itself names.
So run the assessment on yourself first. Then do the sustained work of mitigating the risks you find.
I am not a lawyer, and this paper is not legal guidance. Counsel’s framework does that work well. What follows is the other side of the table: how an organization becomes the one that has the answers, and how its leaders navigate towards a positive conclusion.
In matters of the law
Compressed to its essentials, counsel’s framework begins with scoping decisions: the structure of the transaction, the extent and materiality of AI use, whether the target develops AI or deploys it, and the industries and jurisdictions in play. It then asks five kinds of questions.
Proprietary development. Who owns the algorithms and the training data, what third-party components sit inside them and on what terms, how the models were trained and validated, and whether anyone is watching them for drift now that they are in production.
Third-party AI. How deeply the organization depends on systems it did not build, what it has customized, what its contracts actually permit, and whether it vetted those providers before adopting them.
Deployment and controls. Which uses are internal and which face customers, which fall into high-risk categories, and whether configurations are documented well enough to survive a change of ownership.
Training data. Where it came from, how it was prepared, what licenses and consents govern it, and whether anyone has tested it for bias.
Generative AI. Who owns the outputs, how consistent and accurate they are, and what restrictions apply to their use.
Two further inquiries wrap around these five.
The first concerns the regulatory landscape, from the EU AI Act and the growing patchwork of US state law to the consumer-protection and sector regulators who were never AI regulators but have jurisdiction anyway. This latter concern is often overlooked by organizations, particularly in the United States, who see a lack of AI regulation as justification for their own lack of AI diligence. For example, if an AI system causes an organization to commit financial fraud, it doesn’t matter if the AI is itself not regulated. The crime itself is regulated regardless of whether or not it was committed by AI. This is what Centru calls AI’s second-order regulations: rules that do not govern AI directly, but govern the conduct in which AI takes part.
There is another vector inside the regulatory landscape that few technology leaders have on their radar at all. Counsel’s framework names export controls and trade restrictions. Most CIOs, if they think about jurisdiction, think about data residency and sovereignty. Almost none think about whether the models they depend on can lawfully be made available to their own staff. For example, in June 2026, as Anthropic disclosed at the time, the US Commerce Department directed it to suspend access to two frontier models for any foreign national, inside or outside the United States. Anthropic disabled both for every customer within hours, and access was not restored for nineteen days. To be clear, no customer had done anything wrong. The control landed on the technology, and while the models in question had not been available long enough for most organizations to have come to depend on them, every customer of those models absorbed the disruption nonetheless. Export controls reach AI software, hardware, and model weights. If your architecture assumes a particular model, that assumption belongs on a risk register.
The second of the two wrapping inquiries concerns organizational governance. The policies, training, oversight, incident response, and whether any of it extends to the AI the organization buys rather than builds.
Two observations from within the framework hand the baton to the technology side.
The first is that the authors treat missing documentation as a result rather than a gap. Where a target cannot produce records of how its AI was built, tested, and governed, that absence “should itself be considered a diligence finding.” This is exactly right, and I will build on it.
The second is that the authors are candid about the limits of the data room. Documents describe intent. Interviews with the people who build and operate the AI describe reality, and the framework directs counsel toward those interviews. These investigations cannot be carried out by lawyers alone; rather, their success is the product of partnership between experts in matters of the law and experts in the technology itself.
Read together, the two observations reveal what the checklist actually measures. Every item on it is either an artifact a well-run AI program already produces in the ordinary course of its work, or it is evidence that no such program exists. The buyer’s counsel is not really asking about the AI. They are asking how the organization runs.
If that is the question, answering it takes an instrument. The one I use is Centru, which I co-authored, in its third release as of this writing and published openly by the Center for Trustworthy AI: five pillars, twenty-five weighted dimensions, and a five-level maturity model that measures an organization against them. It was written, in part, with diligence in mind. More than that, it was written to describe what a well-run AI program looks like from the inside, which turns out to be what counsel is trying to establish from the outside. The Center publishes its specification in full so that the reasoning behind any finding can be examined. I will not rehash its mechanics here. Where this argument turns on a model or an instrument, I say enough to follow the argument and no further, because the specification is published in full and stands on its own as the companion to what follows.
Andrew Welch, Christopher Huntingford, Ioana Tanase, and Ana-Maria Welch, Centru 3, Center for Trustworthy AI, 2026, https://centerfortrustworthyai.org/centru.
Four theses
Counsel’s framework tells us what will be asked, but I see that line of investigation as incomplete from the technology and organizational leader’s side of the table. On this I take four positions, none of them a restatement of the legal view. The first two represent complementary halves of the same idea.
Documentation is an asset, not overhead
Counsel’s framework treats absence of documentation as a finding in its own right. I go further. The presence of rigorous documentation and other substantial markers of a mature AI program is worth more than their absence costs, because these markers are not descriptions of the work. They are byproducts of how the work is run.
The markers are few, yet visible. Among them, an executive vision that has been stated rather than assumed, a roadmap carried by the programmatic rigor to advance it, data governance that holds under inspection, and an operating discipline that produces records without being asked for them.
For example, institutions that can produce their workload inventory, data lineage, measurement results, incident and response history, and the record of their governance decisions do not merely document well. They are far more likely to run well, and the artifacts are the residue. Those that cannot produce them have said something about their operations that no interview will walk back. The gap between what an organization claims about its AI and what it can show is where the real finding lives, and it is why the first hours of any diligence are so predictive, for those hours are a live demonstration of chaos or rigor. The artifacts answer before humans can. Evidence that already existed is also worth more than evidence assembled for the occasion. The strangers at the gate can tell the difference.
The materiality inversion
This thesis runs—uncomfortably—in the opposite direction. Conventional reasoning might hold that an organization with minimal AI use warrants a lighter look. There is less to inspect, so inspect less. This is backwards.
AI’s maturity in an organization cannot precede that organization’s use of AI itself. Operational maturity develops only through deployment, and through the discipline that deployment forces on data, governance, and operations. An organization with no meaningful AI in 2026 has not avoided risk. It has simply failed to act. This carries a risk of a different kind, not one created by AI, but one created by its absence. The risk is to the organization’s vitality, to its ability to keep pace with its peers, its competitors, and the constituencies it exists to serve.
This is not a legal risk, so it does not appear on counsel’s checklist. The risk is entirely strategic, going to valuation in a transaction and to the capability to succeed in its particular mission when in steady state. In either case, the treatment of thin AI use is not a narrower scope but a different question. Not “what are they running,” but “what have they failed to do, and what will it cost to do it now?”
There is a harder version of this, and I have watched it happen more than once. A new chief executive or managing director arrives twelve to twenty-four months ahead of an anticipated transaction with a mandate to improve the numbers. Cost comes out. Headcount comes out. Technology is a soft target because its returns are slow and its costs are legible, so the data work is paused, the governance function is left unfilled, the architects who understood the estate are let go. The profit and loss improves. Then the diligence begins, and the buyer’s experts find an organization that cannot describe its own AI, cannot produce its lineage, and has no one left who remembers why anything was built the way it was. The EBITDA looks better. The asset is worth less.
The two absences are two sides of the same coin. Whether the gap is in the records or in the capability, the organization that cannot show its work has told you how it runs... and how it thinks.
The spectrum problem
Regulation, and the diligence patterns built atop it, are largely binary. An organization is a developer of AI or a deployer of it. Its obligations follow from this binary.
For example, a vendor developing AI products is deeply scrutinized for the integrity of the products themselves. A customer deploying AI products built by someone else is scrutinized less on the particulars of those products, but for the conduct of their deployment: its own diligence in product selection, contracts, deployment practices, colleague training and development, etc.
Real organizations increasingly work at neither pole. They rather live along a spectrum, from “embedded AI” that arrives inside tools they already buy, through “extensible AI” built on vendor platforms with their own data and workflows, to “differential AI” engineered in-house to tackle the problems nothing off the shelf can solve. Centru organizes its AI Workloads pillar along exactly this spectrum and observes that the line between extensible and differential has already blurred. The same agent can be characterized either way depending on how deep the customization runs and whose data it reasons over.
Extensibility does not shrink the grey area. Extensibility is the grey area.
It expands that grey area, quickly, and often without anyone explicitly deciding that it should. An organization that has not classified its own workloads along this spectrum will have them classified by someone else, a regulator or a buyer, with no stake in a generous reading and every incentive to apply the developer’s obligations. Where AI-specific law is thin, the second-order regulations reach the spectrum regardless. Sector regulators do not care which side of the developer line you stood on when the harm occurred.
Readiness cannot be manufactured under pressure
The dimensions that carry the most weight in a diligence exercise are the ones that take longest to mature. Data governance. Operational discipline. The digital fluency of the workforce, which in some regimes is now a legal duty rather than good practice. For example, Article 4 of the EU AI Act (Regulation (EU) 2024/1689) requires providers and deployers alike to take measures supporting AI literacy among the staff who operate AI systems on their behalf, an obligation softened but not removed by the 2026 Digital Omnibus amendment (Regulation (EU) 2026/1744). Centru characterizes several of these dimensions as “slogs”, for they pay off across years rather than quarters, and none of them can be meaningfully built out in the weeks between a letter of intent and an open data room. An organization in the tunnel must surface what exists, frame it honestly, and disclose gaps with a credible plan. But it cannot close those gaps. Those on the horizon are offered only a bit more runway. Eighteen months offers enough time to fix the fast-moving dimensions, executive vision and workload prioritization among them, and to begin addressing those that occupy a longer time horizon, but not to finish them. Readiness cannot be manufactured.
This is why steady state is the scenario that matters most, not least because every organization occupies it. It is the only scenario in which meaningful work can actually be done. The city is built in peacetime. By the time the strangers arrive at the gate, the walls are what they are. The question, then, is no longer whether to build them but what to say about the ones that were never put up.
Three scenarios, one target state
Earlier I introduced the three scenarios under which an organization is subject to AI diligence. Each asks something different of the organization, and it is worth being honest that only one of them is generous. Counsel’s framework is principally concerned with organizations “in the tunnel” (Scenario One), and secondarily with those “on the horizon” (Scenario Two). Technology and other institutional leaders do not have the luxury of looking over the short horizon or waiting for the tunnel itself.
Scenario One: In the tunnel
Diligence is weeks away or already underway. A transaction’s letter of intent is signed or nearly so, the data room is being populated, and counsel on the other side has a list. Or worse, should the organization find itself not in a transaction, but in a legal or regulatory-driven investigation, potentially facing civil or criminal charges.
There is no time to mature. There is time to find out how mature you already are, and to say so—to own your destiny as best you can—before someone else figures it out.
Time is of the essence. Run the maturity assessment in days rather than weeks. Prioritize the dimensions that carry the most weight before considering the others. Gather evidence that already exists. Do not manufacture what does not. Interview your own technical experts before those carrying out the diligence can do so. To be clear, they will be interviewed. The only question is whether you hear from them directly (and first), or whether you hear from them via whatever diligence produces. Classify every AI workload along the spectrum so that the developer-deployer question is answered on your terms. Then disclose what you found, gaps included, with a plan and a cost against each. A disclosed gap with a credible plan is priced. A discovered gap is a discount, and it takes the trust in everything else you said with it.
Scenario Two: On the horizon
Diligence is plausible within six to eighteen months. Nothing yet initiated. This is the scenario in which the most is possible per month of effort, and the one most often wasted, because eighteen months feels like plenty until it is six. It is also the window in which the pre-transaction burndown does its damage.
The work, here, is sequencing. Run your full assessment now, and let the weighting tell you where to spend. Centru weights each dimension by its impact against the effort it takes to move, which is precisely the arithmetic of a remediation budget. The dimensions that score high on both are where a year of work changes the answer a buyer’s counsel will get. Fix those first. Take on the higher-effort dimensions knowing that you may not finish before you’re in the tunnel, but that they are nonetheless critical to emerging intact on the other side. And build the body of evidence as the residue of closing real gaps rather than as a project in its own right. Evidence hastily assembled for the occasion reads as precisely what it is.
One trap deserves its own warning. Averaging across dimensions flatters. An organization strong in the visible dimensions, its vision, its roadmap, its embedded tools, can carry a respectable composite score over weaknesses in data governance and operations at which nobody on the inside is looking. A diligence team is trained to look at exactly those. The composite score is for the board. Dimension-by-dimension scores are for the data room.
Scenario Three: Steady state
The further out a transaction may sit, the less likely it is to be anticipated at all. In the public sector, none is possible. It is, therefore, easy to conclude that diligence is someone else’s problem. This is the conclusion I most want to argue against.
Widen your aperture to see that strangers with a list are everywhere. Enterprise customers now send AI questionnaires with major procurements. The questions on them are counsel’s questions written upon different letterhead.
Insurers moved on this before regulators did. Cyber insurance carriers increasingly ask what models you run, under what controls, and with what governance. In the United States, for example, the Insurance Services Office issued three generative AI exclusion endorsements for commercial general liability policies carrying a January 2026 edition date, CG 40 47, CG 40 48, and CG 35 08, which carriers may attach at renewal to strip out losses arising from generative AI. What an organization attests on a renewal form binds it at claim time, which makes the underwriting questionnaire a diligence exercise with a payout attached. I expect this trend to continue, and expand.
Boards carry fiduciary exposure for AI decisions they did not know were being made. Regulators inquire, and in the public sector so do inspectors general, auditors, legislative committees, the press, and—in many open societies—citizens acting under their jurisdiction’s freedom of information regime. None of these actors requires a financial transaction to open a file. Every enterprise sales cycle is a small diligence. Every audit is a rehearsal for a large one.
So the work in steady state is not preparation for diligence. It is the operating condition that makes diligence unremarkable, the assessment run on a cadence, the evidence produced as a byproduct of running well, the classifications kept current, the gaps known and on a roadmap with owners. Which is to say, it is what Centru describes from beginning to end, arriving here by a different road.
One target state. Three on-ramps. The destination does not move. Only the urgency changes, and how much of the road remains in front of you, where course corrections are possible, rather than behind you, where they are not.
Readiness as the operating condition
AI diligence is an assessment of your AI maturity performed by strangers, under time pressure, with something you value at stake. Much of what I have written between the introduction, where that idea first appeared, and this point has been about the strangers. Who they are, what they carry, when they arrive. It is worth ending on the only variable you control, which is what they find.
Evidence is not the core of what they find. Evidence is an expression of reality, flattering or otherwise. The core is the organization that produced it, or failed to. The buyer’s expert, the regulator’s examiner, the insurer’s underwriter, the inspector general, the citizen with a records request—none of them is really asking about the AI. They are asking whether the institution knows what it is doing, whether it can prove what it knows, and whether anyone is accountable for the difference. These are questions about what you operate. About who you are. About how you are led. They have answers on the day the strangers arrive only if they had answers the year before.
The readiness this paper describes is not a posture assumed for the occasion. It is an operating condition. Assessment run on a cadence rather than under duress. The classifications kept current because they are used, not because they might be inspected. Records that exist because that is how the work gets done. Gaps that are known, owned, and on a roadmap such that you found first whatever someone else finds later. None of that is preparation for diligence. It is the definitional picture of today’s well-run institution, diligence-ready as a side effect of being well run.
The entire argument collapses to a single observation.
The organizations that fare best in diligence are not the ones that prepared for it. They are the ones that never needed to.
Cover image: Rembrandt van Rijn, The Syndics of the Drapers' Guild, 1662. Rijksmuseum, Amsterdam. Public domain.

