Verbatim

← Back to blog

June 27, 2026

The Anna Delvey Problem: When AI Answers Are Too Polished to Question

A composed woman in a black dress and dark glasses sits alone at a courtroom table with a glass of water and a closed folder, a court officer standing behind her, evoking Anna Delvey's too-polished-to-question confidence.

You're deep in an AI thread, working on something important. Internal strategy. A client report. Specs for the vibe coding project.

It's not a single answer, it's a long exchange that gets sharper as it goes. And finally you have it. Clear, confident, convincing, with supporting figures and citations.

It sounds right, and you carry it forward. Nothing to check, because nothing about it asked to be checked.

There is a name for the woman who taught a generation what this kind of trust costs. Anna Delvey.

Anna Delvey convinced a lot of careful people she was someone she wasn't, and she did it not by hiding but by being too polished to question. The clothes, the references, the unbroken confidence.

Nobody pulled the receipts until the money was gone, because nothing in the performance gave them a reason to.

There is a word for the part that comes later, the moment you realize you trusted it and never checked whether it held.

Anna Delveyed (v.): to trust an AI response because it reads confident, polished, and authoritative, only to find out later that it doesn't hold up under scrutiny.

We have all been Anna Delveyed before. What's different is the cost.

Here is why the AI version is the harder con. Delvey was one person. One name, one face, one accumulating trail of unpaid bills. She slipped, and eventually someone followed the trail and caught her.

An AI response leaves no trail. There is no single con artist to investigate, no intent to uncover, no tell that repeats often enough to learn. Every response is a fresh performance with a clean record.

You cannot investigate the model the way you investigate a liar, because the model is not lying. It is dutifully applying fallible reasoning to fallible information, fluently. Every time. At volume.

And the failure is never the entire response being wrong. On the contrary, it's mostly sound, with a few issues that don't hold up.

An unsupported claim. A reasoning gap. A citation error. A missing fact or piece of context that would have changed the direction.

The danger is that those issues live inside work that is otherwise right, so nothing about the response flags them.

What sets the damage is not so much how many issues there are. It's the severity of each issue, and how much rides on the decision it feeds.

One load-bearing issue, under a decision that matters, is enough.

Plot severity against stakes and you can see where a response turns dangerous. We call it the Anna Delvey Matrix.

The Anna Delvey Matrix, also known as the Verification Risk Matrix: a heat map plotting issue severity against the value of the decision. Three stages climb the diagonal from the safe green corner to the dangerous crimson one. Contained: the issue stays in your head, low stakes, easy to correct. Shared: it becomes your work in slides, memos, or emails with your name on it, and the people who trust it are not checking. Released: it leaves the building into reports, filings, code, or submissions with real consequences, hard to take back.

The three stages climb from the safe corner to the dangerous one. At every step up, fewer people can still catch the issue, and the cost of missing it multiplies. The check gets cheaper to skip exactly as it gets more ruinous to have skipped.

At the top, a quiet issue stops being embarrassing and starts being expensive.

None of this is hypothetical, and it is not slowing down. The examples run from last year to the past few weeks, and they are the most credible names in the business.

Last year Deloitte delivered a report to the Australian government, a review of the government's welfare compliance system, for $440,000 AUD.

As Fortune and the Australian Financial Review reported, the report invented a quote from a Federal Court judgment and cited academic works that did not exist.

A single researcher at the University of Sydney caught it, weeks after publication, by reading the footnotes. Deloitte issued a corrected version, disclosed it had used a generative AI system, and refunded part of the fee.

Then it kept happening.

EY retracted a published report after an investigation found most of its citations were broken or fabricated, including a reference to a McKinsey report that was never written.

KPMG, this month, pulled its flagship report on agentic AI after researchers found only five of its forty-five citations pointed to real, intact sources, and after named organizations disputed the claims the report made about them.

The law firm Sullivan & Cromwell apologized to a court for a bankruptcy filing that misquoted the U.S. Bankruptcy Code and cited cases that did not support it.

Same failure every time: AI-assisted text, published under a trusted name, the fabrications surfacing only after the work was already out.

We made this point earlier this month in What Part of Your AI Workflow Is Responsible for Disagreement?. These are not a bunch of interns cutting corners.

They are some of the most capable professional-services firms in the world, staffed by experienced people, with every incentive to get it right and reputations that depend on it.

That is the point. The failure was not a lapse in competence. It was a gap in process.

The work was produced faster than it was checked, and the error surfaced not when the text was generated but when it was used, by which point it was already a published report under the firm's name.

Notice what it took to catch these. The Deloitte report was caught by one academic reading the footnotes by hand, weeks after publication.

The case is not isolated. The other errors were found only after publication, usually by targeted review.

Current oversight catches only a fraction of them, which raises a real concern that many similar documents never receive that level of scrutiny.

Volume is only the obvious half. The deeper problem is the trail.

When the con is a person, you can investigate after the fact and the truth eventually surfaces. When the con is a fresh response every time, there is nothing to investigate.

The check cannot come later. It has to come before.

And the moment an agent is the one acting on the response, later disappears entirely. When a person acts on an unchecked response, at least a person read it first. When an agent acts on it, no one did.

The instruction is to stop trusting polish, though that is harder than it sounds. Polish is the exact signal you are trained to trust.

The more authoritative a response looks, the less your instinct to check fires, which is precisely backward.

The fix is not to read more carefully. You cannot out-read the volume, and you cannot investigate a trail that does not exist.

The fix is to make every response defend itself before you act on it, against something it cannot control. An outside check on every claim, when the consequences matter, at the same speed the output is produced.

This is why we built Verbatim. Adversarial review on your actual AI responses, in place, as you work.

Start free with the browser extension. Try it free →

For adversarial review across reports, filings, code, and the systems that act on them, talk to us. Get in touch →