How We Rebuilt a Production AI Data Verification System in Two Days

Sep 28, 2026by Carrie Bennette

How evaluation infrastructure, BM25 retrieval, and reusable agent tools let us replace the core of a production AI verification system in two days.

We rebuilt Auto Data Verification (AutoDV), Weave’s AI system for checking factual claims in regulatory documents against their source evidence, in two days. The original system was accurate, but its fixed-context, single-shot architecture had become increasingly difficult to extend as customers asked for more control over what to review, which sources to use, and how claims should be judged. The replacement we built caught every planted error in our adversarial test set while better supporting regulatory teams to shape the review around their source materials and workflows.

Importantly, the rebuild moved quickly because a year of unglamorous engineering investment had already done most of the work.

The Anatomy of Regulatory Data Verification

Pharmaceutical regulatory submissions contain many thousands of factual claims. A submission is less a collection of documents than a small library, often assembled over the better part of a decade. In the era before electronic submission they were delivered on pallets, by truck. I find that image oddly comforting: the physical mass of the submission conveyed just how much there was to get right. 

Across all of those documents, factual claims must be accurate, appropriately qualified, and traceable to their source evidence. Small errors in a document set of that size can be difficult to spot. A transposed digit, a changed unit, or a value simply taken from the wrong subgroup may look entirely plausible and still be wrong.

Lesson one: Build your AI evaluation harness before you need it

A production AI system is much easier to improve when you already know how to tell whether a change makes it better. In a previous post, we described the quality framework we use to evaluate AI-generated regulatory content, including metrics for faithfulness, instruction adherence, and structural integrity. That measurement infrastructure is the foundation for being able to quickly answer whether a new approach actually moves the needle.

To evaluate AutoDV, we already had a test set spanning four representative areas of regulatory content, including nonclinical toxicology, manufacturing (CMC), clinical study reports, and clinical pharmacology. It was adversarial by design: we replicated mistakes that are common in both human and AI-based authoring systems. Some of these errors were familiar transcription and versioning mishaps: lost precision, changed units, transposed digits, or a value copied from the wrong subgroup, timepoint, or version of a source document. Others targeted AI-assisted drafting failures, such as “a slight increase in ALT was observed in males at 100 mg/kg/day” from a source document becoming simply “ALT increased with treatment” in a summary, where the AI drops important qualifications that materially shape the conclusion and broadens the finding beyond what the evidence supports. 

We injected those errors alongside changes that should not be flagged, such as formatting differences, permitted rounding, and synonym substitutions, because a useful verification system has to distinguish true errors from benign noise, and not simply flag everything. 

Because the benchmark already existed, we could compare the new agentic AutoDV – which can decide when to search for more evidence instead of working from fixed content – directly with the production AutoDV. 

AutoDV Benchmark results: new agentic system vs. production system 

The new system caught more real mistakes – every error in our test set of over 100 unique planned errors – while also producing roughly half as many false alarms. At the same time, the time for the system to review a document doubled and was roughly 20% more expensive to run. For this feature, that was a tradeoff we were willing to make. Shaving a few seconds from an asynchronous process matters less than catching an error that would otherwise reach a regulatory reviewer. 

Most importantly, none of these conclusions required weeks of user testing or arguments about whether the new output “looked better.” We were able to test a new approach, run the evaluation, and make a confident decision thanks to our existing evaluation framework. 

Lesson two: Choose your retrieval strategy for the information need

Auto Data Verification was my first project at Weave, and my first task was the deceptively simple business of finding the evidence at all. The obvious (and fashionable) answer in 2025 was embeddings and semantic search. Semantic search represents text by meaning, allowing the system to retrieve passages that are conceptually similar even when they do not use the same words. For our particular problem, however, it was also the wrong answer.

Semantic similarity is excellent when the question is, “Which passages are about liver enzymes?” Data Verification usually needs to answer something narrower: “Where did this exact factual claim come from?” In regulatory documents, the most informative parts of a claim are often the least semantic: a study identifier, a treatment arm, a timepoint, or an exact value like 4.7 mg/kg. Several passages may discuss the same finding in almost identical language while differing in the one detail that determines whether the claim is correct or not. We needed retrieval that could point to a value and say: here it is, on page 64, in Table 3, for this dose group.

So we moved the reference engine toward lexical retrieval, using BM25, a widely used ranking algorithm for full-text search. Lexical retrieval works differently than semantic search: BM25 ranks passages based largely on the specific words and terms they share with the query, weighting rarer terms more heavily. For tables, we query each cell’s value with additional context from the table titles and row and column headers to give BM25 the exact lexical signals needed to distinguish treatment arms, endpoints, and timepoints. It was an older and considerably less fashionable technique than semantic search, but for our needs it worked better: retrieval became both more accurate and much faster. At the time, it was probably the least glamorous thing I worked on. Nobody demos a search index, let alone from the 1990s. But it quietly became one of the most important pieces of our system.

With fine-grained retrieval in place, we built the original verification workflow on top of it. For each sentence and table cell in a document, Auto Data Verification retrieved a large set of likely source excerpts and passed them to an LLM to verify. The early architecture had an important constraint: the evidence set was fixed. If the relevant source was missing or incomplete from that set, the model could not go looking for more. We compensated by retrieving generously and supplying much more context than most individual claims required, increasing the odds that the necessary evidence was somewhere in the input.

It worked. We shipped it, and customers used the feature successfully for roughly a year.

The broader lesson was not that BM25 is better than embeddings. Semantic search is very powerful when meaning is the signal. In data verification, however, exact lexical details often carry more information than semantic similarity does.

Lesson three: Build AI systems for replacement

The original system did not fail on accuracy. It aged out on extensibility, which is a less dramatic way to die, but no less final. The single-shot architecture constrained two important things in advance: what evidence the model could inspect and the rules it used to judge each claim. As customers wanted more control over scope, sources, and review criteria, each variation required changing the workflow itself rather than simply configuring the system. 

The rebuild replaced that fixed orchestration with an agentic verification loop in which the model can search for and inspect additional evidence as needed. For each claim, the new verification engine starts with the ten source excerpts most likely to contain the answer. If that evidence is sufficient, it returns a decision immediately. If not, it can search for more, inspect additional source material, and continue until it has enough evidence to decide. In our evaluation, 82% of claims were resolved from the pre-seeded evidence; the difficult cases could now be investigated further rather than being forced to make a judgment from an incomplete context window.

What made that replacement fast was that we did not have to replace everything underneath it. The same lexical search capability we had built for the first Auto Data Verification system became one of the agent’s key tools — and, by then, was already powering retrieval for several other AI features across Weave’s platform. In the old architecture, search ran once up front to assemble a large fixed evidence set. In the new one, the agent can call that same capability whenever it needs more evidence.

That is what we mean by building for replacement. The orchestration was disposable; the useful capabilities underneath it were durable. Search and document access could survive one architecture and be recomposed into the next.

The same principle makes the new system easier to adapt. Scope can be narrowed to a table or section, sources can be constrained, and review instructions can vary without requiring a separate pipeline for each case. Models improve, retrieval strategies change, and customer requirements evolve. We increasingly aim to build AI systems so that when one layer ages out, we can replace it without starting over.

The two-day rebuild that took a year

So, did we rebuild Auto Data Verification and get it into production in two days? Technically, yes. But the headline hides most of the work that made those two days possible.

Most of that work happened over the year before the rebuild: building a robust evaluation harness with adversarial test cases, creating clean interfaces between components, and improving document processing and retrieval for factual verification strategies. By the time we were ready to replace the core verification engine, we could change it without rebuilding the product around it, measure the new system against the old one immediately, and ship with confidence that the tradeoffs were worth it.

The lesson is not that production AI systems can be rebuilt in a weekend. It is that the unglamorous work you do beforehand determines whether the next better idea takes two days to ship – or two months.

Contributors: Melissa Morine, Jason Astorquia, Julie Xu, Dani Bergey, Keith Bagchi, Olga Liakhovich

Are you an engineer who wants to help build AI systems with rigorous evaluation, in a meaningful domain, and where going from experiment to production takes days, not months? See our Careers page for open roles.