The Direct Answer

The best question is not whether a company calls itself an “AI company,” uses a large language model, or has raised money under an AI label. Investors should ask what system the company operates, what decision the system influences, what happens when it is wrong, and whether customers would suffer material harm if it stopped working. As of September 28, 2026, AI diligence should therefore combine conventional business analysis with model-specific review rather than treating AI as a separate asset class. A promising demo may still conceal manual review, weak data rights, unpredictable inference costs, or dependence on a third-party model provider. The commercial question remains decisive: does the product create measurable customer value, and can its gross margin survive after compute, moderation, support, and compliance expenses? For private deal flow, the aim is not to produce a universal rulebook. It is to identify the few technical and contractual facts capable of changing valuation, closing certainty, or post-closing ownership risk.

Also worth reading: How Should AI Founders Target Private Investors in 2026? · What is the definitive AI startup due diligence checklist for private market investors in 2026? · How Do Founders and Investors Use AI for Secure Deal Diligence in 2026?

How to Test Whether the AI Claim Is Material

Begin by separating an AI-enabled business from an AI-dependent one. An AI-enabled business may use automated software for recruiting, search, document review, or customer support while its core economics still depend on people, distribution, or industry expertise. An AI-dependent business may be unable to deliver its principal product if model access, proprietary data, evaluation quality, or inference capacity fails. Diligence teams should request the last 12 months of product-level revenue, gross margin, usage, and customer-retention data, then quantify what portion of those results depends on the model. A useful threshold is to identify any workflow that accounts for more than 10% of revenue, more than 20% of gross profit, or more than five material customer relationships. These are screening thresholds, not legal standards, but they direct attention without declaring every model feature material.

The technical review should also distinguish claimed automation from observed performance. Ask for production logs, incident records, evaluation sets, baseline comparisons, and the percentage of outputs accepted, edited, or rejected by staff. If an AI feature saves 20 hours per week, the team should demonstrate the former workflow, actual time saved, and whether the saved work is reflected in revenue, headcount, or service quality. Disclose whether the company has a fallback process and how long manual operation can continue. A business with a reliable fallback may be resilient, while one that promises fully automated service without human review may have an operational fragility that conventional financial statements do not show.

What Technical Diligence Should Actually Cover

Data diligence should begin with provenance and permission, not the size of a stated data set. Investors need to know whether company, customer, employee, medical, financial, public, or scraped data can lawfully be used for training, retrieval, evaluation, and product delivery. The team should map data sources, licensing restrictions, deletion obligations, cross-border transfers, and any personal-information requests already received. A document claiming to contain “millions of records” is not equivalent to millions of usable, rights-cleared records. For sensitive data, the diligence process should request anonymization methods, access controls, retention schedules, and evidence that test environments contain production-like information. The same discipline applies to synthetic data: synthetic records can reduce some privacy exposure, but they may reproduce bias and can create security risks if they reveal the structure or unusual characteristics of the source population.

Model diligence should identify what was built internally, what was fine-tuned, and what remains a third-party service. Teams should test whether performance comes from the claimed model, customer-specific retrieval, workflow design, human reviewers, or a narrow benchmark. Request an evaluation plan with named success metrics, error tolerances, test-set dates, and known failure modes. The plan should be divided between pre-deployment testing and continuous production monitoring, with responsibility assigned to a named owner. Ask how often the model, prompts, data pipeline, or ranking logic changes, and whether material changes require customer notice. Vendors can accelerate deployment, but a vendor dependency also introduces price, availability, terms-of-service, data-use, and model-retirement risk.

Comparing the Main Diligence Approaches

There is no single review method that answers every AI question. A questionnaire is fast and inexpensive but can elicit polished answers rather than operating evidence. A technical audit provides deeper evidence but costs more and may expose confidential data or interrupt engineering teams. A pilot or red-team exercise tests behavior under selected conditions, although it cannot prove safety across every use case. The appropriate method depends on where AI sits in the business and how much damage a failure could cause.

FeatureInternal ReviewExternal Technical DiligenceLive Pilot or Red Team
Typical cost$5,000-$25,000 for a focused internal effort$25,000-$150,000+ for a specialist review$10,000-$75,000+ depending on scope
Time to start1-3 weeks3-8 weeks2-6 weeks
Best evidenceLogs, architecture, interviews, internal testsIndependent architecture and data reviewObserved behavior on controlled tasks
Main weaknessTeam bias and limited independenceCost and access to sensitive systemsNarrow tests may miss rare failures
Best forEarly-stage or lower-risk applicationsMaterial data rights, infrastructure, or model dependenceHigh-impact, customer-facing, or regulated systems
These ranges are practical planning estimates rather than market-wide published prices. A company with proprietary models, regulated data, or safety-critical deployment may require a broader assessment costing more than $150,000. By contrast, a narrow internal review may be enough for a low-stakes feature with no sensitive data. The key is to match spending to potential loss rather than applying the same process to every AI opportunity.

Turning Findings Into Investment and Deal Terms

Diligence findings should affect valuation, structure, covenants, or closing conditions rather than remaining in a technical report. A missing data license may require a purchase-price adjustment, escrow, indemnity, or condition precedent. Dependence on a single model provider may justify protections covering service interruption, data deletion, minimum notice of termination, and a transition period. Weak security controls may require remediation milestones after closing, with reporting rights until completion. If historical AI-generated revenue is materially overstated, the parties should test customer demand without the disputed feature, verify cohort retention, and assess whether discounts or price reductions are needed. Sellers may describe technical capability as an asset, but buyers should price only the capability they can verify and control.

Representations should be specific enough to be enforceable. A broad statement that the company “owns all necessary intellectual property and complies with applicable law” may miss the operational detail behind model training, dataset licenses, open-source software, contractor work, and provider restrictions. Counsel should review whether the disclosure schedules identify material models, datasets, material evaluation results, incidents, and government inquiries. Warranty protection is not a substitute for diligence because collecting a claim after closing can be difficult and disruptive. It is best understood as residual protection for gaps that diligence could not reasonably resolve.

Board and management teams should also assign ongoing ownership after the transaction. A dashboard might track model-drift rates, severe incidents, evaluation pass rates, data-deletion requests, security findings, vendor concentration, and unit inference cost. Thresholds should trigger escalation: for example, more than 5% of outputs manually corrected in a critical workflow, a 20% increase in inference cost per customer, or any unresolved severity-one privacy or security event. These figures should be calibrated to the product, but having a trigger before it is breached is more useful than debating risk after an incident.

Common Mistakes in AI Due Diligence

The first common mistake is allowing a technical label to substitute for a business case. Language such as “proprietary,” “autonomous,” or “real-time” has no consistent meaning. Diligence should translate those terms into measurable claims and ask whether they affect revenue, cost, customer retention, or competitive advantage. A second mistake is reviewing only the model while ignoring the surrounding system. Retrieval, integrations, permissions, human review, monitoring, and incident response often determine production performance more than the model name.

Another error is asking for hundreds of questions without prioritizing them. A long questionnaire creates activity rather than certainty and encourages generic responses. Institutional discussions have warned that more questions and faster answers are not always better; a shorter set tied to material assumptions is usually more defensible. Buyers should prioritize roughly 10-25 core questions, then add targeted follow-ups where evidence is missing. Public reporting also illustrates how AI-generated questions can encode baseless accusations, so every allegation should be labeled as a hypothesis, supported by evidence, and tested against primary documents.

Teams also make the mistake of requesting sensitive data without a controlled review process. Credentials, raw personal information, source code, and customer prompts can create additional legal and cybersecurity exposure. Diligence should use data inventories, redacted samples, secure rooms, read-only access, and written reviewer obligations. Finally, some investors overvalue a polished evaluation score. A benchmark may use training-adjacent data, cover a narrow task, or measure agreement with a reference output rather than business usefulness. Production cohorts, error costs, and customer outcomes remain more persuasive than a single impressive percentage.

When to Pause, Proceed, or Walk Away

A deal should pause when material claims cannot be reconciled across contracts, product analytics, invoices, technical logs, and management explanations. Examples include unexplained training-data rights, inconsistent usage figures, high customer concentration around an AI feature, or a model provider that can terminate service without notice. A pause does not automatically mean rejection; it identifies the condition that must be resolved before money is committed. The parties can narrow the scope, obtain third-party confirmation, secure contractual protection, or reprice the opportunity.

Investors can proceed when the company demonstrates repeatable customer value, a credible path to acceptable unit economics, documented data and intellectual-property rights, and a functioning incident process. Early-stage companies will not have perfect systems, and perfection is not a realistic approval standard. The relevant question is whether known weaknesses are proportionate, disclosed, funded, and improving. A 12-person company with a 6-person product team, for example, cannot be expected to maintain the governance structure of a 1,200-person enterprise, although basic access control, data deletion, backups, and vendor review should already exist.

Walking away is reasonable when the economics depend on unowned data, the product cannot outperform a simpler non-AI alternative, expected compute costs consume the gross margin, or management cannot explain the most important technical failure mode. Legal compliance does not guarantee commercial value, and technical capability does not guarantee transferability. The strongest case is one where the system is not merely impressive in demonstration but remains useful, economical, and contractually durable after normal adverse conditions.

A Practical Diligence Sequence

The first practical step is to establish materiality and reduce access risk. Counsel and the investment team should identify the product, revenue exposure, data categories, decision impact, and applicable regulatory concerns before requesting materials. Within several business days, the company should provide a concise architecture diagram, system inventory, data-flow map, vendor list, model-provider agreements, and product-level performance summary. The reviewer should then compare those materials with customer contracts, privacy notices, security policies, board materials, and general-ledger accounts. Disagreement between documents is often more informative than a missing document because it identifies a control weakness.

Next, the team should run a focused technical session with the founder, product lead, security owner, and data lead. The session should test the production architecture, evaluation design, human escalation, customer support burden, and unit economics over at least the last 12 months. Where justified, an independent specialist should inspect a redacted repository, sample a secure data room, reproduce selected evaluations, or conduct a limited adversarial test. Findings should be graded as deal-stopping, price-affecting, remediable, or monitor. This grading converts technical criticism into an investment decision while avoiding the false choice between accepting every risk and rejecting the company.

The final step is to document who owns each unresolved issue, what evidence would close it, and by what date. For the Mercer Club network, this process is most useful as confidential deal screening: founders can understand the evidence buyers need, while operators can compare approaches before an off-market process becomes constrained. It should not be framed as an automatic endorsement or rejection score. The output is a structured basis for conversation among founders, technical advisers, counsel, and investors, with the private deal-flow network providing context and access rather than replacing professional judgment.