We asked 5 famous AIs
what fans think.
Then we asked the fans.
Claude. GPT-5. Gemini. Grok. We handed every one of them live web search and asked simple questions about a real fanbase — 6,229 Jason Statham fans and the 65,000 comments they wrote on Reddit. Then we checked their answers against what the fans actually said.
Ask an AI twice, get two answers.
We asked Claude the same question three times. It said 15%, then 75%, then 25% — same question, same day, nothing changed. In one pair of identical runs a model swung 47 points and the whole ranking reshuffled. Confident numbers with nothing underneath.
Our engine doesn't guess. It reads the fans' real comments and counts. Run it twice, you get the same answer twice — down to the individual yes and no — with the receipts: how many comments, what they said, which way each leaned.
Half the error of the best AI on earth.
Eight questions about the Statham fanbase. To keep it honest the answer key was graded by the rivals' own AIs — OpenAI, Anthropic and xAI models, majority vote, none from our vendor. Lower = closer to what fans really think.
To rule out tuning-to-the-test, we then wrote fresh questions across two arenas and ran them stone cold — engine frozen, one shot each. On the eight with a solid answer key: first again, 10.6 vs 14.1 for the best AI. Different rivals got lucky in different rounds. Only one was near the top every time.
Every AI said yes — 58% to 85%. It's the obvious answer, it's what the internet says. But when fans name their actual favourites, The Transporter barely comes up.
The obvious answer was wrong. The crowd knew better — and only the one actually listening to the crowd caught it.
We know things Google doesn't.
We went digging in the fans' wider Reddit lives — 79,000 posts beyond the movie threads — for things no search engine can surface. The verdicts are decisive, and every one comes with the quotes behind it.
It reads star ratings out of pure text.
8,324 Letterboxd cinephiles, 192,000 written reviews, 2.3 million star ratings. When we hold the ratings, these questions aren't predictions for you — they're counts, including cross-taste joins no page on the internet has ever published. So to make it a real test, we blindfolded our own engine: reviews only, never the stars — and had it call the hidden numbers from words alone.
Half a million words of “solid”, “glorious trash” and “masterpiece” — turned back into star ratings it was never shown. That's the blindfolded version, and it matters because most audiences have no stars: Reddit threads, YouTube comments, your community are text with no ratings attached. In product mode there's no blindfold — your numbers are counts, exact.
Your audience isn't on Google.
Statham was the hardest possible test for us — world-famous, so the whole internet is his fan data and every AI gets to search it. We won anyway. Now flip it: your customers, your community, your fans aren't searchable. Their real favourites, their inside jokes, what they'd actually buy — none of it is on a page. There, the AIs have nothing to look up, and the engine holding the corpus is the only one left in the room.
Ask an AI, you get a vibe with a percent sign. Ask us, you get a measurement — the number, the comments behind it, and the same answer tomorrow.
Competitors ran via OpenRouter with native web search (:online); Grok also used X search.
Keys — rival-vendor majority judging (Reddit) · pure star-rating arithmetic (Letterboxd).
Every run audited by cmd/dev/benchaudit from raw receipts in data/validation/results/runs/.