Datumplate

How to scope a first AI document automation build

Most AI pilots fail before any code is written. A scoping method for legal and insurance teams that want a working system, not a demo.

Most AI projects fail. The commonly cited figure is around 80 percent, roughly double the failure rate of ordinary software projects. Having watched this category up close for years, we think the number is real, and we think most of those failures were visible on day one, in the scope, before anyone wrote a line of code.

This is the scoping method we use on our own builds. It is written for legal and insurance teams because that is who we build for, but the logic travels.

Start with one workflow, not a platform

The failed pilot usually began as a platform: automate our document review, transform our claims process. A scope like that has no edges, so it can never be finished, only abandoned.

A scope with edges names one workflow, one document type, one data source. Not "automate contract review" but "extract renewal dates and termination clauses from our vendor agreements." Not "AI for claims" but "pull the loss details out of first notice of loss emails into the claims system."

The test: can you describe the workflow in one sentence that names the input, the output, and who touches it today? If the sentence needs an "and," it is two workflows, and the second one goes on a list for later.

Choose the success metric before the vendor

Every failed pilot we have examined shared one trait: nobody had agreed, in writing, what number would prove it worked. Without that number, the pilot ends in a demo, the demo impresses, and then nothing ships, because impressing was never the same thing as working.

Pick one measurable number and its target before anyone builds anything. Good ones are boring: hours of manual review removed per week, extraction accuracy against a hand-checked sample, turnaround time per document. One metric, a current value, a target value. If a vendor resists committing to a metric, that is the cheapest red flag you will ever get.

Demand the risk in writing

Every AI system has a weak point. For document work it is usually the ugly ten percent: the scanned PDF from 2011, the amendment that contradicts the base agreement, the handwriting. A scope that does not name its main risk is not optimistic, it is unexamined.

Ask whoever is building, internal team or vendor, one question: what is most likely to make this harder than it looks, and what is the plan when it does? A specific answer means they have done this before. A confident "no major risks" means you are the risk.

Place the human before you place the model

In regulated work, the question is never whether a human reviews, it is where. Decide the review point as part of the scope: does a person approve every output, a sample, or only the cases the system flags as uncertain? This single decision drives the accuracy the system actually needs, which drives cost more than any other choice. A system that drafts for human approval can ship months before one that must be right on its own.

Prove it small before you buy it big

Whatever the eventual system costs, the first commitment should be small enough to walk away from: one narrow use case, real data, evaluated against the metric you chose, priced fixed so the cost cannot drift. If the result clears the bar, scale it with evidence in hand. If it does not, you have bought the second most valuable outcome available: a cheap, early, documented no.

Be suspicious of anyone who wants to skip this step, and equally suspicious of a pilot that cannot fail. A pilot with no metric, no risk register, and no walk-away point is not a test. It is a sales process wearing a lab coat.

The one-page version

Before any build starts, you should be able to fill one page: the workflow in one sentence, the single success metric with current and target values, the main risk and its mitigation, the human review point, and a fixed price for proving it small. If you cannot fill the page, the project is not ready. If a vendor cannot help you fill it, they are not either.

This page is, not coincidentally, the shape of the free written diagnostic we send teams who ask. If you would rather we filled it in for your workflow, email us and we will reply within five working days. No call required.

All articles