The Direct Answer: Build an AI Diligence System Before a Deal
Founders should build AI diligence governance before a private transaction by creating a documented system for identifying every relevant AI system, classifying its risk, testing material claims, assigning decision rights, and preserving evidence. The system should cover more than the software being sold or invested in. It must include the models, datasets, prompts, integrations, vendors, permissions, human review processes, and downstream decisions that could transfer with the deal.
Also worth reading: What Are Treasury Governance Policies and How Should Founders Evaluate Them in 2026? · What is the definitive AI governance checklist for founders in 2026? · How Should Founders and Operators Use AI for Investment Due Diligence in 2026?
A private transaction can change the risk profile of an AI product in ways that a product demonstration does not reveal. A buyer may receive production data, employee or customer information, regulated records, model weights, API credentials, or obligations to third-party providers. It may also inherit an AI system that performs a more sensitive function after integration into a larger platform. The right question is therefore not whether a company “uses AI” or whether its AI is “responsible.” It is whether a particular system is fit for a defined purpose, under known operating conditions, with controls proportionate to the likelihood and severity of harm.
A founder can begin with a lightweight process, but a lightweight process should still produce durable answers. For example, the company might complete an inventory in 30 days, conduct a preliminary risk review within 14 days of receiving diligence materials, and require senior approval before granting production access to sensitive data. Governance becomes meaningful when another person can reconstruct what the company knew, what it decided, and who accepted the residual risk.
Why AI Diligence Is Different from Ordinary Technology Review
Conventional technology diligence often focuses on source code ownership, architecture, uptime, cybersecurity, scalability, and intellectual property. AI diligence adds questions about training data, model behavior, evaluation quality, human oversight, and the circumstances in which a system can produce harmful or legally consequential outputs. A model can have strong security controls and still be unsuitable for employment, credit, health, education, insurance, or public-benefit decisions.
The distinction matters because AI claims are often probabilistic. A supplier may state that its system is “90% accurate,” but that number may refer to a benchmark that does not resemble the buyer’s data or decision. Accuracy can also hide subgroup disparities, false-negative rates, distribution shifts, or performance degradation after deployment. Before a deal, the buyer should ask what the 90% means: which task, which population, which time period, which threshold, and which costs of error.
Data rights create another layer. A company may possess a technically capable system but lack a defensible basis to use the data, retain derived artifacts, or transfer them to an acquirer. The buyer should distinguish among customer data provided for inference, data used for training, data used for evaluation, synthetic data, public data, licensed data, and internally generated annotations. It should also identify whether a model or embedding can expose or reproduce personal or confidential information.
AI diligence should therefore combine product review, data review, legal review, security review, and operational review. Legal counsel alone cannot determine whether a system is reliable. Engineering cannot assess all licensing, privacy, or employment implications. Governance is the mechanism that connects those disciplines to a single approval decision.
Define Scope Before Reviewing the Technology
The first practical step is to define what counts as an AI system for diligence purposes. A useful definition includes machine-learning models, generative systems, automated decision tools, recommendation engines, speech or image recognition, large language model integrations, retrieval systems, and AI-enabled analytics. It should also capture less visible components, including vendor APIs, data enrichment services, model-monitoring tools, and automated agents that can take actions in external systems.
Scope should follow capability and consequence, not branding. A rule-based spreadsheet may not require the same review as an autonomous agent, but a simple model used to rank loan applications may deserve greater scrutiny than a creative writing tool. Founders should identify systems that can influence people’s access to money, employment, health, education, housing, insurance, legal services, or safety. They should also identify systems that handle confidential information, generate external communications, or execute transactions.
The review should cover the current environment and the proposed future state. An AI tool that is merely experimental today may become a core product feature after the transaction, especially if the buyer expects cross-selling or integration. Conversely, a sophisticated system with limited authority may present less risk than a low-cost chatbot connected to a customer database and able to issue refunds or modify records.
A useful classification can have three levels. Low-impact internal tools might include drafting assistance with no sensitive data. Medium-impact systems might support customer service recommendations or sales prioritization with human approval. High-impact systems might make or materially influence decisions about individuals, access to services, or regulated activities. These categories should influence evidence requirements, approval authorities, and post-deployment monitoring.
Create a Repeatable Review Process
A repeatable process does not need to be bureaucratic. It should, however, contain consistent gates, reviewers, questions, and records. Founders can start with a short intake form that asks for the system’s purpose, owner, users, data sources, model and vendor, deployment method, affected populations, decision authority, monitoring arrangements, and known limitations. The form should require links to relevant policies, contracts, evaluations, incident records, and technical documentation.
The next stage should test claims against evidence. “Bias testing completed” is not enough. The reviewer should know which test was used, when it was run, what dataset it covered, what metrics were reported, who commissioned the work, and whether the results can be reproduced. Claims about security, privacy, accuracy, explainability, or regulatory compliance should be supported by reports, test results, architecture diagrams, access-control evidence, or contractual commitments where appropriate.
A practical process might use four gates. The first gate confirms ownership and scope. The second assesses data, model, security, legal, and operational risk. The third determines whether the system may proceed in a sandbox, pilot, or production environment. The fourth establishes post-approval conditions, including monitoring, human review, incident reporting, and retirement criteria.
The process should also specify who can approve what. A product leader may approve a low-risk internal drafting tool, while a cross-functional committee should approve a system using health or financial data. Legal and security teams should have the ability to stop deployment when evidence is missing, but they should not become permanent bottlenecks. Clear service-level expectations, such as a five-business-day review for a low-risk internal tool, help preserve accountability without creating delay.
The Core Diligence Questions
Founders should ask questions that reveal the system’s real operating boundaries. They should ask what happens when the model is uncertain, when the input is incomplete, or when the user behaves differently from the training population. They should determine whether a human can meaningfully override the system or whether the human is merely asked to approve an output they do not understand. They should also ask whether the system can be used for purposes beyond those originally authorized.
The buyer should test whether the supplier’s claims are independently verifiable. Requesting a demonstration is useful, but demonstrations are curated. Ask for a representative evaluation suite, failure cases, incident history, model-change records, data-provenance documentation, and details about performance under edge conditions. If the system interacts with external tools, inspect permissions, logs, transaction limits, and the consequences of a manipulated prompt or poisoned document.
A useful comparison is between capability and authority. A model may generate a recommendation without automatically making a decision. That is not equivalent to a system that automatically approves a loan, schedules an employee termination, changes a medical recommendation, or sends money to a third party. Governance should be stricter where errors are difficult to reverse, affect vulnerable populations, or expose the company to contractual, regulatory, or reputational liability.
The buyer should also investigate whether the company can separate the AI system from the business if something goes wrong. Are the models and data portable? Can the system be disabled without disrupting all operations? Does the vendor retain broad usage rights? Is there a termination plan? A well-governed acquisition does not assume that every risk can be corrected after closing.
Compare Risk Levels Rather Than Applying One Standard
| System or use case | Primary concern | Evidence before approval | Typical governance response |
|---|---|---|---|
| Internal drafting assistant with no sensitive data | Hallucinations, confidential input, uncontrolled use | Approved use cases, data classification, user guidance, output review | Owner-level approval, acceptable-use terms, periodic sampling |
| Customer-service recommendation tool | Incorrect responses, escalation failure, sensitive records | Representative accuracy tests, escalation rules, access controls, monitoring | Human approval, performance thresholds, complaint and incident reporting |
| Employment or candidate-ranking system | Bias, explainability, disparate impact, employment law | Validated subgroup testing, human review design, audit trail, legal assessment | High-impact review, documented decision rights, ongoing testing and appeals |
| Credit, insurance, or benefits decision support | Accuracy, fairness, regulatory exposure, financial harm | Outcome validation, data provenance, error analysis, model-risk review | Senior committee approval, limits, human override, suspension criteria |
| Autonomous agent with external transactions | Fraud, prompt injection, unauthorized actions, weak auditability | Permission controls, transaction limits, simulation, security testing, kill switch | Restricted pilot, real-time monitoring, dual authorization, immediate shutdown authority |
Founders should resist the temptation to approve a system because its vendor calls it “human in the loop.” Human review is meaningful only when the reviewer has enough time, information, authority, and incentive to intervene. If the reviewer sees hundreds of outputs per hour and cannot understand the system’s confidence, the human is a ceremonial checkpoint rather than a control.
Mistakes That Create False Confidence
One common mistake is treating an AI policy as though it were governance. A policy describes desired conduct, but governance requires proof that the conduct occurred. A signed code of ethics does not show that access to production data was restricted, that an incident was investigated, or that a model was reevaluated after a material change. Policies should therefore be connected to workflows, approvals, logs, and owners.
Another mistake is relying on vendor certifications or broad security questionnaires without examining the actual service. A vendor may have strong enterprise controls but use a third-party model with unclear data retention, training use, or geographic processing terms. The buyer should identify every subprocessor and determine whether the vendor’s commitments flow through to the customer. Contracts should address confidentiality, breach notification, audit rights, model changes, data deletion, return of assets, and assistance with regulatory inquiries.
Diligence also fails when founders ask only what can go wrong but not who will bear the cost. A risk register should assign an accountable owner for each material issue, state the treatment plan, and identify the date by which the issue must be resolved. Risks without owners are warnings, not controls. A risk accepted indefinitely by an unnamed executive is usually an unrecorded assumption.
Finally, companies sometimes conduct extensive review before a deal and then stop monitoring immediately after closing. AI behavior changes when users adapt, data drifts, vendors update models, and integrations expand. The acquisition agreement may assign diligence obligations to a team that no longer exists. A transition plan should name the receiving owner and preserve the right to suspend or reverse a deployment.
When to Act and How Quickly
A founder should establish AI diligence governance before signing a letter of intent, term sheet, or exclusivity agreement whenever AI is material to the company’s value, product roadmap, data assets, or risk profile. If the seller says AI is merely a feature, that is a reason to clarify its role, not a reason to defer. The parties should agree on what technical information will be exchanged, which data may be used in diligence, and whether models, prompts, weights, or evaluation artifacts will be delivered.
Timing should reflect the deal, not an arbitrary annual cycle. For a small seed financing or a $500,000 asset acquisition, a two-week intake and a focused review may be proportionate. For a $50 million acquisition involving healthcare, financial services, or sensitive personal data, the buyer should expect several weeks of technical, legal, security, privacy, and model-risk work. The exact number is less important than setting explicit milestones before confidential information is widely shared.
A staged approach can reduce friction. First, identify systems and high-impact uses. Second, obtain the minimum evidence needed to decide whether to proceed. Third, conduct deeper testing only where the risk justifies it. Fourth, record approval conditions and post-closing obligations. This approach is often better than sending a 100-question questionnaire to every vendor before the buyer understands which systems matter.
Governance should be tightened immediately when an incident, regulatory inquiry, security breach, unexplained performance decline, or model change occurs. It should also be revisited before a system moves from pilot to production, expands to a new population, connects to a new data source, gains authority to take actions, or is transferred to a new legal entity.
What a Founder Should Hand to the Next Owner
The final deliverable is not a polished memo. It is an operating record that the buyer, board, investors, regulators, and future management team can use. That record should identify the company’s AI inventory, systems by risk tier, data and vendor dependencies, material findings, approval decisions, unresolved risks, contractual restrictions, monitoring indicators, incident procedures, and retirement or transition plans.
For a private deal-flow network such as The Mercer Club, the practical value of this work is that it makes AI claims comparable across founders and operators. A seller should be able to explain not only what its system does, but how it knows what it does not know, who controls the consequences, and what evidence supports continued use. That is the standard by which a private transaction can preserve innovation without transferring avoidable risk.
Founders should remember that governance is not a brake applied after commercial excitement. It is a way to make commercial decisions durable. The strongest system may be modest: a named owner, a clear definition of purpose, a few well-chosen tests, explicit limits, and a record of every material decision. What matters is that the organization can state, with evidence, that it knows how the AI system works, where it may fail, who is responsible, and when it must be stopped.