Verbatim

← Back to papers

THE VERIFICATION LAYERPAPER 02

Protection Against Decision Risk

How adversarial review lets verification keep pace with AI generation.

Verbatim2026

Summary

Paper 01 named a gap. AI systems produce work that is fluent, confident, and finished-looking, and they produce it faster than anyone can check it. The result is that a growing share of AI output supports a decision without ever being verified. We called that gap unverified confidence, and a 2025 global study of 48,000 people puts a number on it: two thirds report relying on AI output without evaluating its accuracy.[1]

This paper is about how to close it.

We argue that verification is not a step people should improvise at the end of a workflow. It is a layer, and it belongs in the workflow as deliberately as generation does. We set out what that layer is for, where it lives, and the three things you can adjust when you build one.

The layer we describe is built on adversarial review: putting one AI's work in front of competing models trained by rival labs, and keeping what survives. Three levers control how hard it presses. How many examiners. How capable they are. And how deep the examination goes.

The finding that matters is that these compound. The reference standard for checking anything is a domain expert with the sources open, and few organizations can afford to apply that standard at the volume AI now generates. The question this paper answers is how close a machine process gets. Across five contested questions and more than 1,400 cross-model critiques with nine to eleven models each, we have not observed the result fail. That is depth on a small set of hard cases rather than broad coverage, and we report it as such. At lighter settings the result tracks the configuration, which is why the levers deserve to be understood: they are the difference between a review that protects a decision and one that goes through the motions.

1. The bottleneck, and what it costs

The speed of AI generation has created a bottleneck for human judgment. Where judgment cannot keep up, individuals and organizations carry decision risk they cannot see. AI workflows need a verification layer that runs at the speed of generation.

That is the argument of this paper. The rest of this section is why the bottleneck exists.

An AI system can produce a market analysis, a client memo, a legal summary, or a technical specification in seconds. Reading that output carefully, tracing the claims in it, and deciding whether each one holds takes a person the same amount of time it always did. Possibly longer, because the output arrives polished, and polish makes text harder to scrutinize rather than easier.

So the ratio moved. When one person produced three documents a week, careful review of all three was normal. When the same person produces thirty, review of all thirty is not going to happen. Something gets cut, and what gets cut is the checking, because the checking is the part nobody sees.

This is a bottleneck, and it has a specific shape. The constraint is not intelligence and it is not access to information. It is human attention, applied to verification, at a volume that human attention cannot reach.

Our position is that the answer is not to slow generation down and not to ask people to read more. It is to build verification that runs at the speed of generation, and to make its output something a person can act on in the time they actually have.

There is evidence that the checking is already being cut rather than merely likely to be. The same 2025 global study, run by the University of Melbourne with KPMG across 47 countries, found that 66% of respondents reported relying on AI output without evaluating its accuracy, and 56% reported making mistakes in their work as a result.[1]

2. What a verification layer is, and what it is not

A verification layer is a defined step between generating something and acting on it. Its job is to separate the parts of an output that are defensible from the parts that need a person to adjudicate.

That is the entire claim, and the modesty of it is deliberate.

A verification layer is not an arbiter of truth. It does not tell you what is true. It tells you which parts of a piece of work would survive someone pushing on them, and which parts would not. Anyone claiming more than that for an automated process is describing something that does not exist.

It is not editing. Verification reports on a piece of work. It does not hand you a rewritten version. This distinction matters more than it sounds. If a process returns an improved draft, you can no longer see what was wrong with the original, and the improvement has been taken on trust exactly the way the original was. Findings are auditable. Replacements are not.

Note that findings still drive improvement. In practice, the most common thing to do with a set of findings is hand them back to the model that produced the work and let it refine its answer. The layer does not do the rewriting itself, and that separation is what keeps the loop honest: generation and verification stay two different operations even when they run seconds apart.

It is not the same for every piece of work. How much verification a document earns depends on what happens next, and we return to that in section 4.

One field already built one

The closest analogy is software. None of this is a new idea there: software engineering has run a verification layer for decades and does not call it that.

Type checking, automated tests, continuous integration, code review, staged rollout. Each catches a different class of error, none of them catches everything, and they run continuously and cheaply at roughly the speed code is written. A developer does not decide whether today's work is worth verifying. The layer is simply there, and work that fails it does not proceed.

Software got there first for a structural reason: code is executable, so a large share of its ground truth is machine-checkable. You can run it and see. Prose has no equivalent, which is why the same infrastructure never appeared around memos, analyses, and reports, and why the people producing them have improvised instead.

The field also has a name for what accumulates when the layer is skipped. Ward Cunningham introduced the debt metaphor for software in 1992, to explain to non-technical stakeholders why cleanup work deserved budget, arguing that shipping quickly is a loan rather than a gift and that unrepaid shortcuts charge interest on every subsequent hour of work. The compound term "technical debt" came later; the 1992 report uses only the metaphor.

Decision risk is not the same shape, and the difference is worth keeping. Technical debt is a stock: it sits in an artifact that persists and gets built upon, and its cost compounds. Decision risk is an exposure that resolves at a moment. You act on the output or you do not, and then it either cost you something or it did not. An unverified shortcut in a codebase gets more expensive every month. An unverified claim in a client memo is either load-bearing on the decision or it is inert.

What software offers is not the metaphor. It is the proof that a verification layer is ordinary infrastructure rather than an added burden, and that a field which builds one stops arguing about whether verification is worth the time.

3. What verification actually checks

Before comparing methods, you have to say what they are all looking for. Every verification approach we are aware of, from a person rereading their own draft to an automated system running many models, is producing some version of the same three findings.

What holds. The parts that survive checking.

What is claimed without support. Statements presented as fact with nothing behind them. This is not the same as wrong. It is the category of things the reader would be trusting rather than knowing, which is a different and often larger category than people expect.

What is missing. Considerations not addressed, reasoning not shown, alternatives not weighed. This is the hardest of the three, because absence leaves no trace in the text. Nothing on the page tells you what is not on the page.

The fields change with the content

Those three findings are constant. What fills them changes depending on what is being checked, and this is where most thinking about verification goes wrong.

A research summary is mostly factual claims, so checking it means checking claims against sources. A client memo is not. A memo contains facts, and those facts are traceable, but a memo rarely fails because a figure is off. It fails because the ask is unclear, or the framing is wrong for the person receiving it, or an assumption underneath the argument was never examined. A creative brief fails on different things again.

So the questions change. One live implementation of this uses five field sets, one per content type:

  • Analytical. Verified, disputed, reasoning gaps, recommendations.
  • Business. Audience fit, clarity of ask, weak claims, stronger framings.
  • Technical. Correctness, clarity, edge cases, alternatives.
  • Creative. Voice, specificity, tension, untried angles.
  • Conversational. Tone, intent match, what is missing, better said.

Same three-part shape underneath, different fields on top. Naming a stronger framing is not a rewrite; it is the evidence that the current framing is weak, and it is still a finding about the work rather than a replacement for it.

Applying a fact-checking rubric to a strategy memo produces one of two failures. It flags a deliberate strategic choice as an unsupported claim, which is a category error. Or it returns almost nothing, because there were few checkable facts in the document to begin with, and the emptiness reads as a clean bill of health.

4. What sets how much verification is warranted

Two inputs, and they are usually collapsed into one.

Decision risk. What happens if this is wrong. A draft explored internally and a document filed with a regulator carry different exposure, and it is reasonable for them to receive different amounts of scrutiny. This is the input people already reason about informally.

The nature of the content. This one gets missed. Risk is not only a property of what you do with an output. It is a property of what the output is. Some work is risky in ways no source can resolve. A strategy memo built on a bad assumption is dangerous, and there is no document you can open to catch it. A financial claim is dangerous in a completely different way, and there is a document, and opening it settles the matter.

This means content type does two jobs at once. It sets what you check for, which is section 3. And it sets what kind of risk you are exposed to, which determines how much checking is worth doing and of what type.

Decision risk has two properties, not one

Magnitude is the obvious one: what it costs if this is wrong.

Reversibility is the one that gets left out, and it moves independently. Jeff Bezos framed this in Amazon's 2015 shareholder letter as one-way and two-way doors: some decisions can be walked back cheaply, and some cannot be walked back at all, and treating the second kind like the first is how organizations get hurt.

For verification the consequence is direct. A large decision you can reverse next quarter tolerates thinner checking than a small one you cannot reverse at all, because the cost of being wrong includes the cost of getting back. Irreversibility raises the verification a piece of work earns regardless of its size.

What verification does not reach

Decision risk is broader than anything a verification layer can address, and the paper should be plain about which part it touches.

Verification addresses the information component: whether what you are deciding on is true, supported, and complete. It does nothing about whether the decision itself is wise. It does not tell you the timing is wrong, that a better option existed that nobody proposed, or that the whole framing of the choice is mistaken. A perfectly verified analysis can still lead to a bad decision.

What verification removes is one specific failure: acting on something because it sounded right, when nobody checked. That is a narrower claim than managing decision risk, and it is the one we can make.

5. The three levers

If you are building or buying a verification process, there are three things you can adjust, and together they determine how much of the expert-with-sources standard you reach. Each catches things the others miss, which is why they compound rather than compete. The most common mistake in the field is running one and assuming it covers the rest.

One AI answer passes through adversarial review by rival models and produces a structured critique with four parts, returned to the original model or a person.

Adversarial review returns a structured critique, not a rewrite.

Lever 1: Count

How many independent examiners, meaning rival models, check the work.

One model checking its own output catches slips it can notice on a second reading. It cannot catch anything it is wrong about in a way it cannot see, which is the condition Paper 01 describes.

A second model changes that immediately, because it was trained by a different lab on a different mixture and is wrong about different things. Several models change it further, because their blind spots stop lining up at all. This is the engine of adversarial review: rival systems have no stake in each other's answers and no shared habit of error, so what survives all of them has earned something no single model can grant its own work.

Lever 2: Capability

Which examiners you use.

A capable model reasons better about the work, is more likely to look something up, and is more likely to notice an omission. This is independent of count. A single strong examiner can reach further than several weak ones, and the deepest process staffed by weak models produces a thorough examination of very little.

Capability is not one number, and it is not one price. A model that is excellent on code may be ordinary on regulatory language, and the reverse, so choosing the examiner is part of pulling this lever: the content sets what you check for, and it should also inform who checks. And because intelligence is now rented per request rather than installed, the choice is open every single time. Open-weight models are now capable enough to serve as examiners, often at a fraction of the price, which means a strong panel no longer implies a premium bill.

Whether a strong examiner substitutes for several weaker ones, and at what ratio, we cannot say. That is an empirical question and the answer is likely to depend on the content. We name the lever. We do not publish an exchange rate.

Routing and difference pull against each other

There is a tension inside this lever that is worth naming.

Routing optimizes for fit. It selects the model most likely to be right about this specific material. Examination optimizes for difference. Its value comes from examiners whose blind spots do not overlap, which is why a second model helps at all.

Push routing hard enough and those goals separate. If one model is genuinely best on a class of content, the fit-optimal panel is several instances of it, or several models trained in similar ways to be good at the same thing. That panel is highly capable and highly correlated. It will agree, confidently, and its agreement will mean less than it looks like it means.

The failure mode to watch is therefore not a cheap panel. It is a uniform one. A panel selected purely for capability can be more agreeable and less informative than a mixed panel that includes a model nobody would have chosen on its own.

Lever 3: Depth

How many layers of examination sit on top of the answer.

The shallowest version: examiners check the answer. That is one layer.

Deeper: every response in the process gets examined, not just the original. When multiple models each produce answers and critiques, each model's output is itself cross-examined by the rivals, so nothing in the record stands untested.

Deeper still: the examination runs through a structured sequence rather than a single pass. A turn that specifically demands what is missing. A turn that forces a choice rather than allowing a hedge. A turn where each examiner audits its own prior position. Structure exists because unstructured examination tends to settle into agreement, and agreement is comfortable and not informative.

Depth and count are separate. Twelve examiners each looking once is high count and shallow depth. Three examiners working through structured turns is low count and considerable depth. Neither is a stronger version of the other.

Where sources come in

There is a fourth thing that matters, and it is not a lever. It is a behavior of the examiners themselves.

An AI system can back a claim in two entirely different ways. It can retrieve a document, read the relevant passage, and cite it. Or it can produce the citation from memory, the way a well-read person produces a half-remembered reference in conversation. Both come out looking identical on the page. Recalled citations are frequently correct, which is precisely what makes the wrong ones dangerous: when a recalled citation fails, it fails quietly, and nothing in the output distinguishes it from one that was fetched and read.

Whether an examiner searches or answers from memory is a function of the model, made claim by claim. It is not a dial anyone turns. What the levers do is shape the odds: a capable examiner goes and looks when a claim warrants it, and a larger panel raises the chance that at least one examiner opens the source. We log every search conducted in every adversarial review pass, so for any run we can see how much of the examination was grounded in retrieved documents rather than in memory.

6. What the layer reaches

Model errors are substantially decorrelated in practice. Different systems are trained on different mixtures, filtered differently, and tuned differently, so one system's mistake is unlikely to be another's, an effect the multi-agent debate literature we cite documents directly. That decorrelation is why adding examiners works, and the levers harvest it aggressively. They catch a large share of what is there, which is why a well-configured panel catches so much more than people expect from a process with no human in it.

What remains is the error models hold in common. Models are trained on overlapping material, and where that material is wrong or incomplete, every model drawing on it inherits the same gap. Agreement among examiners measures how correlated they are. It does not measure whether they are right.

This is why searching examiners matter structurally. Examination among rivals clears a large share of the error that is not shared, and an examiner that opens a source does something else on top: it brings in information that was not inside any of the models. In our runs the two happen together. The honest description of a fully engaged layer is examination that generates its own evidence.

It is also why agreement should be reported as agreement rather than as verification. Both are useful findings. They are not the same finding, and most systems present them identically.

7. Calibrating to decision risk

The levers exist to be set, not to be maxed. This is where section 4 and section 5 meet.

Where decision risk is high, press every lever. Many examiners, the most capable available, examination in structured layers. Where decision risk is lower, you have room to choose which levers to pull and how hard.

Level 1: Single rival review

Four settings of the same method. Each card shows how much decision risk it is built to carry.

The economics make this an easy call, and it is worth being direct about them. A single capable examiner reviewing an AI response typically costs pennies, not dollars. A full panel at maximum depth costs somewhat more. The decisions this protects are worth orders of magnitude more, and the gap between those numbers is the entire case for the layer. Verification tuned to decision risk is not an expense to be managed. It behaves like insurance in the workflow.

The one guideline that survives every configuration: when trimming, cut count before capability. Fewer strong examiners beat many weak ones, because a strong examiner reasons better, notices more, and goes to the source when a claim warrants it. A panel of weak examiners produces a thorough examination of very little, and its agreement is the least informative kind.

The failure mode to avoid is tuning everything for factual tracing. A strategy memo resting on an unexamined assumption has no document that settles it. What catches that failure is examination with depth: an examiner willing to challenge the premise rather than the figures, and a structure that forces the challenge rather than allowing agreement. A verification process tuned entirely for fact-checking will return a clean result on a memo whose central assumption is wrong, and the cleanliness will be sincere.

8. Where the layer lives

A verification layer is not one product. It is a capability that shows up on different surfaces, and the surface determines what happens with the findings.

In place, as you work. The review runs where the content is generated. You are reading an AI's answer, you invoke a rival model or a panel against it, and the findings appear beside the work. What usually happens next is the refine loop: the findings go back to the original model, which revises its answer with the challenge in hand. Generate, pressure-test, refine, in the same window, in seconds. This is the everyday surface, and it is where the bottleneck from section 1 actually gets relieved, because verification is happening at the same speed and in the same place as generation.

The full-pressure record. The same methodology at its maximum setting, run on a fixed question with a full panel of rival models across structured turns, recorded claim by claim. Every response cross-examined by every rival at every turn, with the synthesis mapping where the rivals agree and where they diverge. Run publicly, this surface is proof of methodology that anyone can inspect. Run inside an organization, it is the deepest form of the record described next.

The document worklist. For long-form artifacts carrying serious stakes, a legal brief, an enterprise report, a regulatory filing, the layer can decompose the document into individual claims and trace each one to a source or flag it for attention. The output here is different in kind: not feedback for a model to refine, but a worklist for a person to adjudicate. The human acts as a judge ruling on flagged items rather than an editor rereading the whole document, which is what makes review possible at volume. This is the newest of the three surfaces and the least standardized anywhere in the industry.

The three surfaces share one methodology and one discipline. The findings are always separate from the work, so a reader can always see what was challenged and what survived.

9. Everything leaves a record

If section 8 is where the layer appears, section 9 is what it leaves behind. Every adversarial pass is recorded, and that record may matter to an organization more than the review itself.

What was examined. Which rival models examined it. What they flagged, what survived, what was refined in response. Which searches were run along the way, so grounded findings can be told apart from remembered ones. All of it tracked, all of it pointing back to the source it examined.

Here is why that matters. Decisions get questioned later. When one does, the difference between "we were careful with the AI" and evidence of care is a record. An organization that ran its AI output through independent rival models before acting on it, and can show the findings, has evidence that the output was checked before it was relied on. An organization that cannot produce that record has an assertion.

This is what auditability means for a verification layer, and it is the part regulators, boards, and opposing counsel are increasingly likely to ask about. Not whether the answer was right, which nobody can promise, but whether anyone checked, how hard, and what they found. The record answers exactly that question, claim by claim, after the fact, without relying on anyone's memory of having been careful.

The record proves process, not truth. That is consistent with everything in section 2, and it is enough: diligence is a process property, and the record is what makes diligence demonstrable rather than claimed.

References

Examination of one model's output by another has a direct antecedent in LM vs LM: Detecting Factual Errors via Cross Examination (Cohen, Hamri, Geva, Globerson, EMNLP 2023, pages 12621 to 12640, aclanthology.org/2023.emnlp-main.778, arXiv:2305.13281). Their examiner questions the model that produced the claim, using self-inconsistency as an error signal. What we describe is broader: rival models do the examining, panels replace the single examiner, every response in the process is itself examined, and the questions asked need not be factual.

Panels of models examining a question across multiple rounds are treated in Improving Factuality and Reasoning in Language Models through Multiagent Debate (Du, Li, Torralba, Tenenbaum, Mordatch, ICML 2024, pages 11733 to 11763, arXiv:2305.14325) and in ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate (Chan, Chen, Su, Yu, Xue, Zhang, Fu, Liu, ICLR 2024, arXiv:2308.07201). Claim-level tracing against retrieved evidence has antecedents in RARR (Gao et al., ACL 2023, pages 16477 to 16508, arXiv:2210.08726) and FacTool (Chern et al., arXiv:2307.13528, 2023).

The debt metaphor for software originates in Ward Cunningham, The WyCash Portfolio Management System, OOPSLA '92 Experience Report, SIGPLAN OOPS Messenger 4(2), pages 29 to 30, doi 10.1145/157710.157715, available at c2.com/doc/oopsla92.html. The report uses the metaphor throughout but not the compound term now standard in the field.

The one-way and two-way door framing of reversibility is from Amazon's 2015 letter to shareholders, where the two are also labeled Type 1 and Type 2 decisions.

[1] Gillespie, N., Lockey, S., Ward, T., Macdade, A., and Hassed, G. (2025), Trust, attitudes and use of artificial intelligence: A global study 2025, The University of Melbourne and KPMG, DOI 10.26188/28822919. Surveyed over 48,000 people across 47 countries, November 2024 to January 2025. The figures cited are the global ones; per-country figures differ, and the same reliance measure reads 57% in the study's Australia reporting.

If your organization checks AI output before acting on it and wants to test its own method against this one on a shared task, write to us. That comparison is the next paper.