Every Data Lab study starts as an idea pasted into a Suprmind thread, where five frontier AIs try to take it apart. The few that survive get scoped to real data, briefed, and run through Research Symphony. Three made it out so far: 1,324 production turns scored, 178 model releases tracked and six hallucination benchmarks cross-referenced.
7 days free. No credit card. The trial runs GPT, Claude, Gemini and Grok.
All five AIs, Debate and Red Team come with Pro at $45/mo.
The lab is a symbiotic team. Human researchers pick the questions, set the methods and sign off on what gets published. AI researchers do most of the heavy lifting: five frontier models arguing inside Suprmind, plus whatever outside AI tools a study needs, from the deep research agents each lab ships to analysis runs over leaderboard datasets.
Most of the thinking happens in Suprmind itself. Which question is worth asking, which data can answer it, what gets added and how to add it. All of it is argued out in shared threads, where every model reads what the others said before it answers.
The monthly bill for all those AI subscriptions, API keys and research platforms is responsible for most of Radomir’s sleepless nights. Radomir Basta, our founder and CEO, insists it is worth it. He insists quietly, usually around 2 a.m.
It usually starts with Radomir. He arrives with an idea that is bold, flashy and, by his own estimate, unlike anything the industry has seen. He pastes it into a Suprmind thread and asks five frontier AIs what they think.
They tell him. One finds the study that already covered it. One points out that the data to prove it does not exist. One asks what exactly would be measured, and how. By the end of the thread the idea is either smaller and sharper, or gone. Most are gone. That is why there are three Data Lab reports and not thirty.
An idea that survives five harsh but fair critics is worth weeks of data work. An idea that does not was going to fail anyway, only later and at a higher price.
It is the same reason Suprmind exists. AI that agrees with you is pleasant. AI that tells you why you are wrong is useful.
Three studies, three questions: how often frontier AIs contradict each other in real work, how often they make things up, and whether new models actually beat the ones they replace. Every report publishes its method, and every number is there to cite.
Original research
April 2026 Edition – The Confidence Trap
Suprmind’s own production data. 1,324 multi-AI turns from 299 users over 45 days, scored for contradiction, correction and unique insight per provider. Where five frontier AIs disagree, who catches whom, and how often a confident answer fails peer review.
51.3%
Gemini’s confident answers contradicted or corrected
9.77×
Perplexity vs Gemini catch ratio
72.1%
Disagreement on financial questions
Live benchmark
Living reference, revised as benchmarks publish
Seven hallucination benchmarks in one report: Vectara HHEM, AA-Omniscience, FACTS, HalluHard, CJR Citation, SimpleQA and SimpleQA Verified. Cross-referenced model by model and read alongside Suprmind’s own multi-model findings.
61 / 202
Models with a positive AA-Omniscience Index
72.6%
Claude Fable 5.1 wrong instead of refusing
73-86%
Drop with web search on, in OpenAI’s own evals
Live index
October 2026 Edition, updated every two weeks
Every premium AI model from 15 labs since 2023, with the day the public could first use it and a blind-vote check on whether it beat the model it replaced. 178 releases, 78 version pairs measured on LMArena and nine live charts you can embed.
4.5 days
Between premium model releases in 2026
56%
Of measured 2026 releases beat the model they replaced
1 in 6
Measured 2026 releases that lost to their predecessor
The Multi-Model AI Divergence Index measures what happens when several frontier models answer the same question in the same thread. Over 45 days, 1,324 production turns from 299 users, with two to five models per turn, ran through Suprmind’s Disagreement/Correction Index classifier, which scored each provider for confidence, contradiction, correction and unique insight.
The headline finding named the report. Confidence and accuracy are different signals. Gemini’s high-confidence answers were contradicted or corrected by another model 51.3% of the time, the highest of the five. When the stakes rose, Claude’s contradicted-while-confident rate fell 7.5 points. Gemini’s barely moved.
What it does not measure: ground truth. The classifier tracks where models diverge, not which one is right. Divergence tells you where to look. Evidence tells you who was right.
Every major AI model hallucinates, and every benchmark measures it differently. Vectara tests summarization. AA-Omniscience tests whether a model answers when it should refuse. FACTS, HalluHard, CJR Citation, SimpleQA and SimpleQA Verified each catch a different failure. Read one leaderboard and you see one slice.
The AI Hallucination Rates & Benchmarks report puts them side by side: 202 models on the AA-Omniscience board, frontier models profiled one by one, 50+ cited sources and a cross-benchmark table revised as new results publish. The numbers disagree with each other, and the report explains why that matters more than any single ranking.
The AI Models Index tracks every premium model released by 15 leading labs since 2023. Each release gets two answers: the first day the public could actually use it, and whether people preferred it over the version it replaced in blind head-to-head votes on LMArena.
It is also the study where the lab put its own tools on trial. The same research brief went to five deep research AIs, and each built a release history on its own. None got it right alone. Two called a 2026 slowdown that never happened. Every disagreement was settled against lab announcements, API changelogs and LMArena’s published data.
“A premium AI model now ships every four and a half days, but this year only a little over half of them beat the version they replaced in blind votes, and one in six lost. The newest model is a candidate, not an upgrade. Test it on your own work before you switch.”
Radomir Basta, founder and CEO of Suprmind
Every report publishes its definitions and method in full, so you can check how each number was made.
Several labs we measure supply the models Suprmind runs. Their weak results get published the same as their strong ones.
The Divergence Index does not measure ground truth. LMArena preference is not task performance. Each report says what it cannot claim.
AI tools gather and draft. Lab announcements, changelogs and benchmark datasets decide what is true.
When Claude scored five deep research runs, including its own, the report said so. The scoring rules were fixed before any scoring started.
AI Models Index charts are licensed CC BY 4.0. The Divergence Index ships 12 aggregate CSVs under the same license. Turn-level data stays private.
Data Lab research is written to be reused. AI Models Index charts and the Divergence Index dataset are licensed CC BY 4.0, so you can use them, including commercially, with credit to Suprmind Data Lab and a link to the source report. For interviews, press kits and data requests, write to [email protected].
Suprmind Data Lab, report name and edition, with a link to the report page. Credit LMArena and Artificial Analysis for their own figures.
Every AI Models Index chart has an Embed button. Embedded charts update with each edition, so your page never quotes stale numbers.
The Divergence Index publishes 12 aggregate CSVs under CC BY 4.0. The AI Models Index dataset is not published, but custom cuts are available on request.
Fact sheet, light and dark charts, the race-for-#1 video and logos. Request it through the press kit form at the end of the AI Models Index.
Red Team and Debate do the ego trimming. Sequential builds the analysis. First Principles rebuilds the question when the obvious framing fails. Research Symphony runs the full pipeline. Switch modes mid-thread and every model keeps the context.
AIs respond one after another. Each reads everything before it. The default and the deepest.
Best for:
Complex analysis, research, architecture decisions
All five respond simultaneously. A sixth AI synthesizes one unified answer with consensus and divergence mapped.
Best for:
Quick decisions, fact verification, time-sensitive calls
AIs argue assigned positions in sequence. Rebuttals and counter-arguments. Minority views preserved.
Best for:
Strategy validation, thesis stress-testing
AIs attack your plan from six angles in sequence: financial, technical, reputational, regulatory, operational, edge cases.
Best for:
Pre-launch validation, risk assessment, investment pre-mortems
Automated research pipeline that retrieves sources, analyses, fact-checks, challenges, and synthesises. Produces 10,000+ word reports with citations.
Best for:
Deep research, comprehensive reports
Strips a question to its fundamentals. Each model names its assumptions, identifies the underlying axioms, then rebuilds the analysis from the ground up.
Best for:
Highest-stakes decisions where convention is suspect
Sequential, Debate, Red Team, and First Principles all use sequential orchestration – each AI builds on what came before. Super Mind mode runs in parallel with a synthesis layer. Chain any combination mid-conversation.
Beyond the Lab
Research is one job the five AIs do. Here are four more, and every output is a real document you can export and send.
Strategy Consultants
Walk into the partner meeting with five frontier AIs already disagreeing on your behalf. Each fabrication caught before slides leave your laptop.
Verdict
Do not acquire at $42M. Revisit at $26M with NRR turnaround proof.
Founders & Operators
Run a $79 vs $149 split through Debate mode. Watch Claude argue retention, Grok argue elasticity, Perplexity ground both in 2026 benchmarks.
Retention curve flattens past $99. The $50 of headroom buys you Frontier-buyer signaling.
Elasticity at this stage is brutal. You’ll lose 31% of conversions for ~22% revenue lift.
2026 SaaS prosumer benchmarks: 38% of $99+ tools see >40% trial-to-paid lift after price reduction.
AI Power Users
Cancel ChatGPT Pro, Claude Pro, Perplexity Pro, Gemini Advanced. One conversation. Five models. Shared context. $95/mo all-in.
Suprmind Frontier
All five models · one thread · shared context
$95
Investment Analysts
Five knowledge bases reference the same question. Build the strongest case for and against before capital gets committed.
“5 AIs were a go-to resource in setting up our new business venture in NYC. From red teaming the initial idea (with harsh feedback), studio market and competitors analysis, to day to day brainstorming about launch phases and website setup. Being able to bounce any idea off 5 AIs, get a clear filtered answer and a todo list in 10 minutes helps a lot.”
CEO, OFF Studio NYC & Funduck Production
“I started using it for competitor research and it just kept expanding – new markets, risk reviews, compliance docs. Five different angles on the same question catches things I would have missed.”
CEO & Co-founder, Miss Amara
“We run everything through Suprmind now – new business ideas, client contracts, marketing strategies. Having five AIs push back on each other in one thread replaced hours of second-guessing between tools.”
Co-founder & COO, Global Digital Marketing Agency
“For analyzing business plans and evaluating client processes, the depth you get from five models reading each other is genuinely different. The Master Document export with custom prompt alone saves me hours on final reports.”
Senior International Adviser, EBRD – European Bank for Reconstruction and Development
Disagreement is the feature.
Bring the plan you are most sure about. Five frontier AIs will read it, argue about it and tell you what breaks, before your money, your client or your board finds out.
7 days free. No credit card required. The trial runs GPT, Claude, Gemini and Grok. All five AIs, Debate and Red Team come with Pro at $45/mo.
FAQ
Suprmind Data Lab is the research arm of Suprmind. It publishes original data on how AI models behave: the Multi-Model AI Divergence Index, the AI Hallucination Rates & Benchmarks report and the AI Models Index. Every study is planned and argued out inside Suprmind, where five frontier AIs challenge the idea before any data work starts.
Both. Human researchers choose the questions, set the methods and review every report before it ships. AI does most of the legwork: five frontier models debate each study in Suprmind, and outside tools such as deep research agents gather and draft. When an AI and a primary source disagree, the source wins.
Yes, with credit. AI Models Index charts are licensed CC BY 4.0, commercial use included, and come with embed code that updates with each new edition. The Divergence Index publishes its aggregate data under the same license. Credit Suprmind Data Lab and link to the report you quote.
Partly. The Multi-Model AI Divergence Index publishes 12 aggregate CSVs under CC BY 4.0, while turn-level data stays private to protect users. The AI Models Index dataset is not published, but you can request a custom cut at [email protected].
It depends on the report. The AI Models Index ships a new edition every two weeks. The hallucination benchmarks report is revised as major benchmarks publish new results. The Divergence Index is an edition-based study built from Suprmind production data, and each edition carries its sample window and date.
It is a fair question, and the reports disclose it. Suprmind runs models from several of the labs we measure, so every report publishes its method, keeps weak numbers for providers in our own roster and names what the study cannot claim. The Divergence Index, for example, measures divergence, not ground truth.
Every research statistic on this page comes from one of the three Data Lab reports linked above, and each report states its sample, window and date. Third-party figures, such as LMArena ratings or AA-Omniscience scores, are credited to their original source inside the report that uses them.
Yes. The lab works in the same multi-AI platform you can try. Spark includes Sequential and Super Mind with a 7-day free trial and no credit card. Pro adds Debate, Red Team and First Principles, the modes the lab uses to trim ideas, and Research Symphony runs on Enterprise.
Disagreement is the feature.
Original AI research, argued over by five AIs first.