The Short Answer: Run a Blind, Scored Test on Your Own Deal
An AI diligence tool bake-off should be treated as an operational experiment, not a software demonstration. As of September 25, 2026, there is no universally accepted ranking of AI tools for M&A diligence because the correct choice depends on the documents, questions, workflow, data controls, and risk tolerance of a specific deal. A strong bake-off begins with 25 to 50 real, permissioned questions drawn from the target company’s actual diligence requests. It then measures answer quality, source traceability, time saved, reviewer disagreement, and failure behavior rather than relying on polished product tours. The winner is the tool that produces the most defensible work for your team at an acceptable cost, not necessarily the system that generates the longest or most fluent answer. For a private deal-flow network such as The Mercer Club, the important comparison is also whether the tool improves the diligence behind a founder or operator opportunity without exposing confidential deal information.
Also worth reading: How Should Founders and Operators Use AI for Investment Due Diligence in 2026? · How Do AI Data Room Review Tools Transform Due Diligence for Private Equity and M&A in 2026? · What Are the Best Practices for AI Deal Due Diligence in 2026?
A useful rule is to allocate roughly 70% of the evaluation to evidence quality and workflow fit, 20% to security and administration, and 10% to presentation. That weighting is an operating recommendation, not an industry standard. It recognizes that a convincing interface is cheap to build, while reliable retrieval, stable citations, access controls, and predictable failure modes are harder. The bake-off should produce a repeatable scorecard that a deal team can defend after the vendor meeting ends. It should also leave an audit trail showing which model, version, prompt, source, and reviewer produced each conclusion.
What an AI Diligence Bake-Off Is Actually Testing
AI due-diligence systems typically perform four tasks: retrieve relevant documents, summarize or extract requested facts, compare documents, and draft analyst notes. Some also identify inconsistencies, generate timeline summaries, map risks, or answer natural-language questions across a data room. These are different functions, and a platform can be excellent at one while performing poorly at another. For example, a general enterprise assistant may answer a question quickly but omit a contradictory sentence buried in a later schedule. A dedicated transaction platform may trace that sentence to an exact page but offer less flexible analysis. Testing a single “write a summary” prompt cannot establish which system is dependable across these tasks.
The test corpus should reflect the real work rather than a vendor’s preferred use case. Include an operating agreement, financial statements, customer contracts, an IP schedule, an employment or equity plan, and internal board materials, subject to appropriate permissions. Add deliberately messy inputs such as scanned PDFs, handwritten notes, spreadsheets with merged cells, conflicting dates, and files uploaded under ambiguous names. A 50-document corpus is not enough to simulate a large data room, but it can expose basic retrieval and citation failures. For a larger process, a 500-document sample is more representative; for an early pilot, 25 well-chosen documents may be enough if every claim can still be checked.
How to Design a Fair, Blind Test
Create the question set before speaking with vendors, and keep each vendor’s identity hidden from reviewers where practical. The same 30 to 50 prompts, source files, and scoring rules should be used for every participating tool. Include at least 20% questions with known answers and another 20% designed to test whether the system admits that the record is insufficient. A useful sample for a mid-sized transaction is 40 questions: 15 fact-retrieval prompts, 10 cross-document comparison prompts, 5 exception-detection prompts, 5 synthesis prompts, and 5 unanswerable or permission-restricted prompts. Those categories should be adjusted to the deal, but the mixture prevents a vendor from winning through one narrow strength.
Freeze the source set during the bake-off so later uploads cannot make one tool appear better simply because it has access to more material. Require each system to cite the file, page or section, and relevant date for every factual assertion. Reviewers should then verify citations against the original documents rather than accepting citation presence as proof. Run each test at least twice because retrieval systems may vary when a large file set is processed repeatedly. A measured pass rate should be recorded at the question level: if a system answers 24 of 40 questions accurately, its observed pass rate is 60%, even if all 24 answers sound confident.
Recommended Scorecard and Vendor Comparison
Score each response from 1 to 5, with half-point increments permitted, across factual accuracy, completeness, citation quality, handling of conflicts, refusal behavior, usability, and administrator control. Factual accuracy should cover whether every material statement is supported by the supplied record. Completeness measures whether the answer addresses all parts of the question and flags material exceptions, while citation quality requires a reviewer to find the supporting passage quickly. Conflict handling should reward a tool that presents competing versions of a fact rather than silently selecting one. Refusal behavior matters because an honest “the documents do not answer this” is safer than a fabricated completion in a transaction.
| Feature | General AI Workspace | Dedicated Deal or Diligence Platform | Human-Led Review |
|---|---|---|---|
| Typical strength | Flexible drafting and broad knowledge | Controlled retrieval, comparison, and transaction workflows | Contextual judgment and accountability |
| Evidence requirement | Depends on configuration and user discipline | Usually emphasizes source links and document-level review | Reviewer must locate and record evidence |
| Best test | Citation accuracy on uploaded deal materials | Cross-document exception detection and permission handling | Interpretation of ambiguous or strategic issues |
| Approximate entry cost in 2026 | About $20-$30 per user per month for common business tiers | Often custom-priced through sales; no reliable universal public rate | Internal reviewer time plus specialist fees when needed |
| Common failure | Fluent unsupported generalization | Expensive configuration or a rigid question model | Slow review and inconsistent documentation |
| Likely fit for a small deal | Yes, with strict review | Possible, but customization may outweigh benefit | Yes |
Test Citation Reliability, Not Just Answer Fluency
An answer that sounds authoritative but points to the wrong clause is a worse failure than a plainly incomplete answer. Require the system to quote or closely paraphrase the supporting passage and identify where it appears. For numerical findings, ask the tool to preserve units, currency, reporting period, and whether a figure is actual, forecast, or adjusted. For contractual findings, request the exact party names, effective date, termination language, renewal terms, and any defined terms that could change the interpretation. A citation labeled only “financial model” is not enough when the requested answer depends on a specific cell or footnote.
Measure citation validity in two ways: whether the cited source exists and whether it actually supports the claim. On a 40-question test, a 90% source-presence rate can still conceal a 20% support error rate if the system cites a real document after drawing the wrong conclusion from it. Therefore, score source presence and substantive support separately. Record unsupported claims as errors even if the final number happens to be right, because a system can reach the correct answer through an invalid chain of reasoning. This is especially important for revenue quality, customer concentration, change-of-control clauses, intellectual-property ownership, and litigation exposure.
Citation quality should also be tested under contradiction. Supply two documents that describe different headcounts, payment dates, or contractual rights and ask the tool to reconcile them. The preferred response identifies the conflict, quotes both records, notes relevant dates, and explains what additional evidence is needed. It should not invent a reconciliation unsupported by the documents. NIST’s AI Risk Management Framework 1.0, published in January 2023, provides a useful governance vocabulary around validity, reliability, transparency, and risk management, although it is not a product certification or an M&A checklist.
Security and Data Handling Are Part of the Test
The security review begins with where files are stored, how they are used for model training, and who can retrieve them. Ask the vendor for its data-processing terms, retention schedule, subprocessor list, encryption approach, user-authentication options, and incident-response process. For confidential deal information, enterprise plans with contractual controls are generally more appropriate than consumer or unapproved free accounts, even if a small business tier appears to support uploads. Do not place client or target information into a system merely because it offers an “enterprise” label; confirm that the intended account and contract provide the required controls.
A practical pilot may begin with synthetic files, public filings, or information that both parties have authorized the deal team to share. Expand access only after access controls, deletion procedures, and permitted user groups are tested. A useful acceptance threshold is 100% compliance with the deal team’s restricted-access rules, because one unauthorized disclosure cannot be offset by good answer scores. For the 2026 bake-off, also test disabled-user access, shared-link expiry, export controls, and deletion from recently used prompts or conversation history. If a vendor cannot explain these behaviors in writing, treat the gap as a procurement issue rather than a minor product detail.
The EU AI Act entered into force on August 1, 2024, with obligations for general-purpose AI models becoming applicable on August 2, 2025 and a broader set of provisions scheduled to apply on August 2, 2026. Not every M&A analysis triggers the same legal requirements, and using an AI tool does not automatically make its output legally binding. Still, buyers should understand their obligations, particularly when AI systems operate within regulated products, employment decisions, credit decisions, or other high-risk contexts. Legal and compliance specialists should determine applicability rather than relying on a vendor’s marketing description.
Practical Steps for Running the Bake-Off
First, appoint one evaluation owner and two or more reviewers, ideally including someone from deal execution, finance, legal, or operations. The owner should prepare the corpus, standardize prompts, record results, and prevent vendors from receiving feedback that changes the test midway. Reviewers should work independently before discussing scores. That reduces the chance that a senior person’s first opinion anchors the rest of the group. Hold a calibration round using one example response from each tool, then revise the rubric while definitions are still consistent across vendors.
Second, ask every vendor to complete a scripted 60- to 90-minute workflow using the common data set. Measure setup time, time to first usable answer, correction time, and total analyst minutes rather than response latency alone. A system that takes 30 seconds to answer but requires 15 minutes to verify its citations may save less time than one that takes 90 seconds and provides a precise page reference. Record login, administrator setup, file processing, prompt formulation, review, and rework as separate steps. On a 40-question pilot, these measurements usually reveal more than a general demonstration does.
Third, conduct a reference check with an existing customer who uses the platform for a comparable workflow. Ask how long implementation took, which promised integrations actually work, how often users must correct outputs, and whether the vendor responded to a security or quality issue. Verify the claim where possible rather than accepting a named logo as independent evidence. Also ask the vendor for model and retrieval configuration details relevant to the test, including knowledge-cutoff treatment, document chunking where disclosed, search settings, and how source updates affect prior answers. Vendors may decline to disclose every architectural detail, but they should be able to explain the behavior a deal team will depend on.
Common Mistakes That Distort the Results
The most common mistake is allowing vendors to choose different questions. One tool may receive a short contract summary, while another receives a complex cross-document analysis, making the scores impossible to compare. Another error is counting a citation as correct without opening it. Demo environments can also distort the result if they contain preselected, clean documents rather than the full set of files your team must process. Give every vendor the same corpus or, if data permissions prevent that, provide parallel corpora with matched content and equivalent difficulty.
A second mistake is confusing novelty with accuracy. A system may produce a more fluent explanation while missing a payment date or misreading a defined term. Fluency is still relevant to user experience, but it should carry less weight than factual support in a diligence process. Do not add AI-generated conclusions to a final memo without naming the source and reviewer. A defensible process is “tool draft, human verification, reviewer approval,” with the reviewer’s name and date attached. The tool should disappear from the approval chain; it should not become an anonymous authority between the analyst and the deal team.
The third mistake is testing only the happy path. Include missing documents, duplicate versions, scanned pages, and questions that cannot be answered from the data room. Test whether the system surfaces limitations or quietly fills gaps with background knowledge. Also avoid selecting a winner before collecting results, because a predetermined preference can turn a bake-off into a procurement ritual. If two systems finish within 5% of each other on total score, choose on narrower operational grounds such as security, existing integration, or reviewer preference. A close score is a reason to clarify the remaining differences, not a reason to manufacture a large gap.
When to Pilot, Buy, or Use Alternatives
A pilot is usually justified when the team repeatedly handles multiple data rooms per quarter, spends meaningful analyst time searching documents, or expects to process a large volume of agreements. A smaller deal team may gain more from well-structured prompts, folder naming, and disciplined human review than from a costly implementation. As a rough starting point, a paid pilot of 30 to 60 days is reasonable if the team can measure a baseline, such as an average of 6 hours spent on document retrieval for each target. The pilot should be judged against that baseline and against a target of cutting review time by 20% without increasing factual errors; these are management thresholds, not published benchmarks.
Alternatives include hiring a transaction analyst, using document-management search, or purchasing specialist review for a narrow issue such as customer contracts or change-of-control provisions. These options can be better when the question depends on local law, an unusual industry structure, or confidential information that cannot leave the organization. A general AI workspace with strict source-checking can be sufficient for early screening, while a dedicated platform may be justified for a larger data room. The right decision is proportional to deal size, workload, and the cost of a missed exception.
Pricing in 2026 should be treated as indicative because vendors change plans and negotiate by volume. Common business AI tiers may fall around $20-$30 per user per month, while dedicated transaction and enterprise deployments are often custom-priced. Implementation, storage, retrieval, connectors, security review, and training can cost more than the headline subscription. The Mercer Club angle is therefore not that every member needs the most feature-rich tool, but that founders and operators can benefit from a structured way to ask what evidence is missing before committing attention or capital. The first deliverable should be better diligence questions, not a more impressive AI presentation.
The Final Recommendation: Buy Evidence, Repeatability, and Control
Choose the tool that makes verification faster, shows its sources, handles missing information honestly, and fits the team’s approved data environment. For a small organization, that may mean a general enterprise workspace with locked permissions and a strict citation policy. For a frequently active deal team, it may mean a dedicated diligence platform that supports document comparison, audit trails, and administrator controls. In either case, keep a human decision owner, record the model and source context, and re-run important findings when the underlying documents change. AI should accelerate the first pass; it should not replace professional responsibility for the final conclusion.
Set a 60- to 90-day review checkpoint after implementation and compare error rate, reviewer time, adoption, and unresolved exceptions with the original baseline. Stop or renegotiate the arrangement if it produces unsupported answers, cannot meet access requirements, or saves less time than the manual process. Do not treat a vendor’s claim of accuracy as a substitute for your own measurements. The most durable diligence advantage is not an AI-generated summary, but a repeatable process in which every important statement is linked to evidence, uncertainty is visible, and people know exactly who verified the result.
For a private deal-flow network, the best question is not “Which AI tool wins?” but “Which tool helps our members make better decisions about the opportunities in front of them?” That framing keeps the evaluation tied to founder and operator value rather than novelty. It also creates a useful feedback loop: recurring questions from members can become a tested prompt set, while documented failures become part of procurement diligence. The result is less theater, fewer unsupported claims, and a clearer path from information to informed action.