The Blind Taste Test Problem: Why Comparing AI Answers Isn't Checking Them

Back in April, LinkedIn set up a Pepsi Challenge of sorts. This time the contestants were AI models.
Here is how it works. You type a prompt. It goes to two models at once, names hidden. You read both responses and pick the one you prefer.
Only then does the blindfold come off.
LinkedIn calls it Crosscheck. It lives on LinkedIn Labs, gated to premium members. The votes get compiled into leaderboards showing which models perform best in specific sectors. Marketing. Engineering. Legal. Finance. Real feedback from professionals in those fields.
Stripped of anything personally identifiable, that conversational data goes back to the AI developers to help them improve their models.
It is not the only tool that lets you compare responses from several models at once. But LinkedIn's version is fun, and hiding the logos does something real. It strips out your bias toward a brand.
At the time, the announcement got a lot of applause. One comment was more cautionary.
"This is useful but it measures output quality, not decision risk." In most B2B environments, the failure isn't that the model gives a 'bad' answer. It's that a 'good-looking' answer gets trusted when it shouldn't. Krista Mollion, in a comment under the launch post.

She's right. And it takes us straight back to the Cola Wars.
In the 1970s and 80s, the Pepsi Challenge put two unmarked cups in front of people and asked which they preferred after a sip of each. Pepsi kept winning.
Pepsi was sweeter, and in a single sip, sweeter had the advantage.
The commercials worked so well that Coca-Cola panicked. They ran the same test internally and got the same result. So they reformulated. New Coke, April 1985, sweeter than the original.
Here is the part everyone forgets. New Coke won the taste test. Coke's own research had it beating both Pepsi and the original formula.
Seventy-nine days later, Coca-Cola brought the original formula back.
The test measured the sip. But people were buying the can.
AI answers have a sweetness too. Fluency. Authority. Polish. The unbroken confidence of a response that never signals doubt.
Put two of them side by side with nothing to check either one against, and sweetness is what you are picking.
Crosscheck blinds you to the brand. It does not blind you to the sugar.
So where is the test that catches the relevant fact the response left out? Or checks whether the claims and citations are accurate? Or finds the place the reasoning goes thin and the argument breaks?
There isn't one.
Hence, the real question. What does it cost you to trust the sweeter response?
There is a second problem, and it is simpler.
The models get updated. Constantly. Grok 4.5 lands and half the received wisdom about rankings goes stale overnight. Most people are still working from a ranking they formed months ago, on versions that have been superseded.
A leaderboard built on last quarter's preferences is a snapshot of a lineup that has already been replaced. Crosscheck now lets you choose which two models to pit against each other, which helps, but it does not change what is being measured.

So what do you actually do about it?
Stop asking which response you like, and start asking how defensible it is.
One simple approach is to take the response you already have and put it in front of a model that did not write it. Different training data. Different instincts. Different blind spots. And no stake in defending the answer.
What holds up? What is weak? What is missing? What would make it better?
That output is not another contestant. It is an assessment of the answer already in your hands, which you can take back to the original model to refine its response.
So by all means, use a comparison tool to pick the model that gives you the sweeter response.
Just know that a defensible answer does not come from picking. It comes from pressure-testing.
Before your AI response goes into a board deck, or a client memo, or a filing, you'll want the answer that isn't just written better, but also survives scrutiny.
That is why we built Verbatim.
Explore the Index.
A single model's output is unreliable in ways that we cannot detect from inside that output. The Verbatim Index positions up to a dozen frontier models against one another, and makes each one defend its answer to the others across four structured turns. The transcript is public: every claim, who challenged it, what held. Read one and watch a confident answer come apart. See how the 2008 financial crisis question held up →
Follow the research.
Models agreeing with each other is not the same as an answer being checked. We are working out what checking actually requires, and publishing it as we go. Read the papers →
Try it on your own work.
The Verbatim Chrome extension pressure-tests an AI response in place, as you read it, using models that did not write it. You end up knowing which parts held and which did not, before you act on any of it. Try it free →
For adversarial review across reports, filings, code, and the systems that act on them, talk to us. Get in touch →