---
title: Suprmind Data Lab
description: "Suprmind Data Lab publishes original AI research: the Divergence Index, AI hallucination benchmarks and the AI Models Index. Methods, charts and data to cite."
url: "https://suprmind.ai/hub/data-lab/"
published: "2026-10-06T16:38:37+00:00"
modified: "2026-10-07T19:58:02+00:00"
author: Radomir Basta
type: page
schema: WebPage
language: en-US
site_name: Suprmind
---

# Suprmind Data Lab

![Suprmind Data Lab: Multi AI platform for decision intelligence and validation.](https://suprmind.ai/hub/wp-content/uploads/2026/10/suprmind-data-lab_suprmind.webp)

> Suprmind Data Lab is the research arm of Suprmind: human researchers and five frontier AIs that argue every idea out in a shared thread before any data work starts. Three reports so far: the Multi-Model AI Divergence Index, the AI hallucination benchmarks report and the AI Models Index.

[Skip to content](#content)





 Original AI Research, Made in a Five-AI Thread


# Suprmind Data Lab: Original AI Research That Survived Five Harsh but Fair AIs**Every Data Lab study starts as an idea pasted into a Suprmind thread, where five frontier AIs try to take it apart.**The few that survive get scoped to real data, briefed, and run through Research Symphony. Three made it out so far: 1,324 production turns scored, 178 model releases tracked and six hallucination benchmarks cross-referenced.



- Grok
- Perplexity
- Claude
- ChatGPT
- Gemini



 [Let the AIs Take Your Idea Apart →](https://suprmind.ai/signup/spark)
 [See Pricing](https://suprmind.ai/hub/pricing/)















 Simplified · Sequential mode
 5 models active


























 ChatGPT
 leans yes



Surface read says yes. The target could double our addressable market, and $42M equals 6x its stated ARR. On the headline numbers, the strategic case looks plausible.
















 Claude
 flag



The retention figure breaks the headline case. At 78% net revenue retention, the existing customer base contracts every year, so a 6x ARR price assumes growth the current economics do not support.
















 Perplexity
 correction



The precedent cited in the target's deck cannot be verified in the supplied files, so it should not anchor the valuation. The price case now depends on retention recovery and credible cross-sell evidence.
















 Gemini
 revised



Revising my read. Once the unverified precedent is removed and retention is priced correctly, the bull case loses its anchor. At $42M, the acquisition does not clear a reasonable diligence threshold.
















 Grok
 pushes back



Pushing back on the pile-on: weak retention does not make the asset worthless. A lower entry price, founder-retention earn-out and evidence of recoverable churn could preserve the strategic case. But those conditions need contractual proof.



























 ChatGPT
 re-priced



Updating the model. Pricing the retention risk and integration cost puts risk-adjusted fair value between $24M and $28M, well below $42M. Inside that band, the strategic upside begins to compensate for the downside.
















 Claude
 conditions



Two gates before close, whatever the price: NRR above 95% for two consecutive quarters, and an earn-out tied to team retention rather than headline revenue. Without both, a discount only buys the same problem cheaper.
















 Perplexity
 evidence



Using the scenario assumptions, a 3x to 4x ARR range is more defensible until retention recovers. On $7M ARR, that supports an offer between $21M and $28M rather than the requested $42M.
















 Gemini
 consensus



Converging: an entry near $26M sits inside the supported range. With the retention covenants and earn-out written into the deal, it survives the main objections raised in this thread.
















 Grok
 risk



One risk remains unpriced: the cost of walking. A rival acquisition could close the strategic window, so the board should compare that downside with the value destroyed by overpaying.












Master Document - Verdict


Do not acquire at $42M. Re-engage near $26M, conditional on two quarters of NRR recovery and a retention earn-out in writing. Otherwise, walk.










Type @ to mention one AI…




























7 days free. No credit card. The trial runs GPT, Claude, Gemini and Grok.
All five AIs, Debate and Red Team come with Pro at $45/mo.








Who Runs the Lab



## A small team of human researchers. A much larger team of AIs.







The lab is a symbiotic team. Human researchers pick the questions, set the methods and sign off on what gets published. AI researchers do most of the heavy lifting: five frontier models arguing inside Suprmind, plus whatever outside AI tools a study needs, from the deep research agents each lab ships to analysis runs over leaderboard datasets.



Most of the thinking happens in [Suprmind](https://suprmind.ai/) itself. Which question is worth asking, which data can answer it, what gets added and how to add it. All of it is argued out in shared threads, where every model reads what the others said before it answers.



The monthly bill for all those AI subscriptions, API keys and research platforms is responsible for most of Radomir’s sleepless nights. [Radomir Basta, our founder and CEO](https://suprmind.ai/hub/about-radomir-basta/), insists it is worth it. He insists quietly, usually around 2 a.m.










Human researchers


Set the question


Pick what is worth measuring, define the method, and check claims against primary sources before anything ships.






Five frontier AIs


Argue it out


GPT, Claude, Gemini, Grok and Perplexity in one Suprmind thread, challenging the premise, the data plan and the conclusions.










Outside AI tools


Do the legwork


Deep research agents, data tools and analysis runs, used wherever a study needs them. One study alone used five deep research AIs.






Primary sources


Settle the facts


Lab announcements, API changelogs, leaderboard datasets and benchmark papers. When an AI and a source disagree, the source wins.














How a Data Lab Paper Gets Made



### Most research ideas die in the first thread. That is the process working.



It usually starts with Radomir. He arrives with an idea that is bold, flashy and, by his own estimate, unlike anything the industry has seen. He pastes it into a Suprmind thread and asks five frontier AIs what they think.



They tell him. One finds the study that already covered it. One points out that the data to prove it does not exist. One asks what exactly would be measured, and how. By the end of the thread the idea is either smaller and sharper, or gone. Most are gone. That is why there are three Data Lab reports and not thirty.





#### The ego trimming is on purpose.



An idea that survives five harsh but fair critics is worth weeks of data work. An idea that does not was going to fail anyway, only later and at a higher price.



It is the same reason Suprmind exists. AI that agrees with you is pleasant. AI that tells you why you are wrong is useful.





- Ideas argued out in shared Suprmind threads before any data work
- Red Team and Debate modes aimed at the premise, not the polish
- Data segments chosen to answer the question, not to decorate it
- Research plans and briefs generated from the thread itself
- [Research Symphony](https://suprmind.ai/hub/modes/research-symphony/) for retrieval, analysis, fact-check, challenge and synthesis
- Humans review every report and check claims against primary sources







 1
 The Idea Arrives
 Usually Radomir’s

Bold, flashy and often late at night. Pasted into a Suprmind thread before anyone can talk him out of it.





 2
 Five AIs Push Back
 Ego Trimming

Red Team and Debate test novelty, feasibility and data. Most ideas end here, politely.





 3
 The Survivor Gets Focused
 Scope and Data

What is left gets narrowed to specific data segments that can actually answer the question.





 4
 Plan and Briefs
 Written From the Thread

The conversation becomes a research plan and briefs, built from what the models settled.





 5
 Research Symphony Runs
 Then Humans Check

Retrieval, analysis, fact-check, challenge and synthesis. Then humans check the claims against primary sources.










The Research



## AI statistics we measured ourselves. Not copied from a press release.



Three studies, three questions: how often frontier AIs contradict each other in real work, how often they make things up, and whether new models actually beat the ones they replace. Every report publishes its method, and every number is there to cite.











 [001





 Original research


### Multi-Model AI Divergence Index

 April 2026 Edition – The Confidence Trap
 Suprmind’s own production data. 1,324 multi-AI turns from 299 users over 45 days, scored for contradiction, correction and unique insight per provider. Where five frontier AIs disagree, who catches whom, and how often a confident answer fails peer review.


 51.3%
 Gemini’s confident answers contradicted or corrected


 9.77×
 Perplexity vs Gemini catch ratio


 72.1%
 Disagreement on financial questions




 Published: April 2026
 Sample: 1,324 production turns
 Users: 299
 Window: 45 days
 License: CC BY 4.0, 12 CSVs


 Read the research ↗](https://suprmind.ai/hub/multi-model-ai-divergence-index/)

 [002





 Live benchmark


### AI Hallucination Rates & Benchmarks

 Living reference, revised as benchmarks publish
 Seven hallucination benchmarks in one report: Vectara HHEM, AA-Omniscience, FACTS, HalluHard, CJR Citation, SimpleQA and SimpleQA Verified. Cross-referenced model by model and read alongside Suprmind’s own multi-model findings.


 61 / 202
 Models with a positive AA-Omniscience Index


 72.6%
 Claude Fable 5.1 wrong instead of refusing


 73-86%
 Drop with web search on, in OpenAI’s own evals




 Benchmarks: 7
 Models: 202 on AA-Omniscience
 Sources: 50+
 Latest revision: October 2026
 Format: Open access


 Read the research ↗](https://suprmind.ai/hub/ai-hallucination-rates-and-benchmarks/)

 [003





 Live index


### AI Models Index

 October 2026 Edition, updated every two weeks
 Every premium AI model from 15 labs since 2023, with the day the public could first use it and a blind-vote check on whether it beat the model it replaced. 178 releases, 78 version pairs measured on LMArena and nine live charts you can embed.


 4.5 days
 Between premium model releases in 2026


 56%
 Of measured 2026 releases beat the model they replaced


 1 in 6
 Measured 2026 releases that lost to their predecessor




 Releases: 178
 Labs: 15
 Refresh: Every two weeks
 Charts: CC BY 4.0
 Embeds: 9 live charts


 Read the research ↗](https://suprmind.ai/hub/ai-models-index/)











Research 001 – The Confidence Trap



## A confident AI answer is not the same as a correct one.







The Multi-Model AI Divergence Index measures what happens when several frontier models answer the same question in the same thread. Over 45 days, 1,324 production turns from 299 users, with two to five models per turn, ran through Suprmind’s Disagreement/Correction Index classifier, which scored each provider for confidence, contradiction, correction and unique insight.



The headline finding named the report. Confidence and accuracy are different signals. Gemini’s high-confidence answers were contradicted or corrected by another model 51.3% of the time, the highest of the five. When the stakes rose, Claude’s contradicted-while-confident rate fell 7.5 points. Gemini’s barely moved.






![Bar chart of AI catch ratios per provider: Perplexity 2.54, Claude 2.25, Grok 0.72, GPT 0.38, Gemini 0.26](https://suprmind.ai/hub/wp-content/uploads/2026/04/chart_2_catch_ratio_asymmetry_preview.svg)*Above 1.0, a provider catches more errors than it gets caught making. Perplexity and Claude made 60.7% of the corrections attributable to a specific model.*Peer corrections


1,401


Logged across 1,324 turns, about one per turn. Each one a claim another model refused to let stand.






Unique insights


3,484


Points one model raised that no other model in the thread did. Five models, five different blind spots.










Turns with signal


99.1%


Turns that surfaced at least one contradiction, correction or unique insight. Only 0.9% stayed silent.






Money beats law


72.1%


Disagreement on financial questions, against 41.2% on legal ones. The models fight 1.75× more over money.**What it does not measure: ground truth.**The classifier tracks where models diverge, not which one is right. Divergence tells you where to look. Evidence tells you who was right.





 [Read The Confidence Trap →](https://suprmind.ai/hub/multi-model-ai-divergence-index/)









Research 002 – AI Hallucination Rates



## Six hallucination benchmarks. One report. No spin.







Every major AI model hallucinates, and every benchmark measures it differently. Vectara tests summarization. AA-Omniscience tests whether a model answers when it should refuse. FACTS, HalluHard, CJR Citation, SimpleQA and SimpleQA Verified each catch a different failure. Read one leaderboard and you see one slice.



The AI Hallucination Rates & Benchmarks report puts them side by side: 202 models on the AA-Omniscience board, frontier models profiled one by one, 50+ cited sources and a cross-benchmark table revised as new results publish. The numbers disagree with each other, and the report explains why that matters more than any single ranking.






![Screenshot of the Suprmind cross-benchmark AI hallucination rates table, August 2026 revision](https://suprmind.ai/hub/wp-content/uploads/2026/10/ai-hallucination-rates-mega-table-aug-2026.webp)*The cross-benchmark reference table, August 2026 revision. The live table covers every frontier model in the report.*Reliability


61 / 202


Models whose correct answers outnumber their wrong ones on AA-Omniscience. The other 141 guess more than they know.






Web search effect


73-86%


Drop in hallucination when web search is switched on, across OpenAI’s own factuality evaluations.










Lowest rate


0.6%


Best score on Vectara’s original summarization dataset. Low on one test does not mean low on all seven.






One update


38 pts


How far Gemini 3.1 Pro cut its AA-Omniscience hallucination rate in a single release, from 88% to 50%.









 [Open the hallucination benchmarks report →](https://suprmind.ai/hub/ai-hallucination-rates-and-benchmarks/)









Research 003 – AI Models Index



## New AI models ship faster. They improve less.







The AI Models Index tracks every premium model released by 15 leading labs since 2023. Each release gets two answers: the first day the public could actually use it, and whether people preferred it over the version it replaced in blind head-to-head votes on LMArena.



It is also the study where the lab put its own tools on trial. The same research brief went to five deep research AIs, and each built a release history on its own. None got it right alone. Two called a 2026 slowdown that never happened. Every disagreement was settled against lab announcements, API changelogs and LMArena’s published data.*Live chart from the AI Models Index. It updates with every edition, and the Embed button on the index gives you the code for your own site.*Release pace


4.5 days


Between premium model releases across the industry in 2026, down from 7.5 days in 2023.






Improvement rate


56%


Of measured 2026 releases beat the model they replaced in blind votes. In 2025 it was 74%.










Regression rate


1 in 6


Measured 2026 releases that lost to the model they replaced by a statistically real margin.






Smaller steps


+9.3


Median LMArena gain per 2026 release, down from +21.5 in 2025.











“A premium AI model now ships every four and a half days, but this year only a little over half of them beat the version they replaced in blind votes, and one in six lost. The newest model is a candidate, not an upgrade. Test it on your own work before you switch.”



Radomir Basta, founder and CEO of Suprmind





 [Explore the AI Models Index →](https://suprmind.ai/hub/ai-models-index/)









How We Keep Ourselves Honest



## Rules the lab follows, even when the data is awkward for us.









### Methods in the open



Every report publishes its definitions and method in full, so you can check how each number was made.







### Bad numbers stay in



Several labs we measure supply the models Suprmind runs. Their weak results get published the same as their strong ones.







### Limits named up front



The Divergence Index does not measure ground truth. LMArena preference is not task performance. Each report says what it cannot claim.







### Sources settle it



AI tools gather and draft. Lab announcements, changelogs and benchmark datasets decide what is true.







### Disclosures on the page



When Claude scored five deep research runs, including its own, the report said so. The scoring rules were fixed before any scoring started.







### Open by default



AI Models Index charts are licensed CC BY 4.0. The Divergence Index ships 12 aggregate CSVs under the same license. Turn-level data stays private.












For Journalists and Researchers



## Cite it, embed it, or ask us for a custom cut.





Data Lab research is written to be reused. AI Models Index charts and the Divergence Index dataset are licensed CC BY 4.0, so you can use them, including commercially, with credit to Suprmind Data Lab and a link to the source report. For interviews, press kits and data requests, write to [[email protected]](/cdn-cgi/l/email-protection#e59597809696a596909597888c8b81cb848c).







### Cite



Suprmind Data Lab, report name and edition, with a link to the report page. Credit LMArena and Artificial Analysis for their own figures.







### Embed live charts



Every AI Models Index chart has an Embed button. Embedded charts update with each edition, so your page never quotes stale numbers.







### Download data



The Divergence Index publishes 12 aggregate CSVs under CC BY 4.0. The AI Models Index dataset is not published, but custom cuts are available on request.







### Press kit



Fact sheet, light and dark charts, the race-for-#1 video and logos. Request it through the press kit form at the end of the AI Models Index.










The Lab’s Toolkit




Red Team and Debate do the ego trimming. Sequential builds the analysis. First Principles rebuilds the question when the obvious framing fails. Research Symphony runs the full pipeline. Switch modes mid-thread and every model keeps the context.































### Sequential

 Default






AIs respond one after another. Each reads everything before it. The default and the deepest.





Best for:



Complex analysis, research, architecture decisions



 [Learn more →](https://suprmind.ai/hub/modes/sequential-mode/)



















### Super Mind

 Fastest






All five respond simultaneously. A sixth AI synthesizes one unified answer with consensus and divergence mapped.





Best for:



Quick decisions, fact verification, time-sensitive calls



 [Learn more →](https://suprmind.ai/hub/modes/super-mind/)



















### Debate







AIs argue assigned positions in sequence. Rebuttals and counter-arguments. Minority views preserved.





Best for:



Strategy validation, thesis stress-testing



 [Learn more →](https://suprmind.ai/hub/modes/super-mind-debate-modes/)



















### Red Team







AIs attack your plan from six angles in sequence: financial, technical, reputational, regulatory, operational, edge cases.





Best for:



Pre-launch validation, risk assessment, investment pre-mortems



 [Learn more →](https://suprmind.ai/hub/modes/red-team-mode/)



















### Research Symphony

 Enterprise






Automated research pipeline that retrieves sources, analyses, fact-checks, challenges, and synthesises. Produces 10,000+ word reports with citations.





Best for:



Deep research, comprehensive reports



 [Learn more →](https://suprmind.ai/hub/modes/research-symphony/)



















### First Principles

 Pro+






Strips a question to its fundamentals. Each model names its assumptions, identifies the underlying axioms, then rebuilds the analysis from the ground up.





Best for:



Highest-stakes decisions where convention is suspect














Sequential, Debate, Red Team, and First Principles all use sequential orchestration – each AI builds on what came before. Super Mind mode runs in parallel with a synthesis layer. Chain any combination mid-conversation.









Beyond the Lab




Research is one job the five AIs do. Here are four more, and every output is a real document you can export and send.


















Strategy Consultants



### M&A pre-mortem in 90 minutes



Walk into the partner meeting with five frontier AIs already disagreeing on your behalf. Each fabrication caught before slides leave your laptop.








 Master Document – preview
 v4 · exported as PDF




#### Skybridge Acquisition – Recommendation Memo



Prepared by Suprmind · Sequential mode · 5 models · 47 min





Verdict



Do not acquire at $42M. Revisit at $26M with NRR turnaround proof.






Executive summary


Five-model consensus matrix


Disagreements & unresolved questions


Risk register (red team output)


Supporting evidence – citations














Founders & Operators



### Pricing experiment, defended



Run a $79 vs $149 split through Debate mode. Watch Claude argue retention, Grok argue elasticity, Perplexity ground both in 2026 benchmarks.






 Debate transcript – preview







 Claude
 PRO – $149




Retention curve flattens past $99. The $50 of headroom buys you Frontier-buyer signaling.








 Grok
 CON – $79




Elasticity at this stage is brutal. You’ll lose 31% of conversions for ~22% revenue lift.








 Perplexity
 CONTEXT




2026 SaaS prosumer benchmarks: 38% of $99+ tools see >40% trial-to-paid lift after price reduction.
















AI Power Users



### Stop reconciling five tabs



Cancel ChatGPT Pro, Claude Pro, Perplexity Pro, Gemini Advanced. One conversation. Five models. Shared context. $95/mo all-in.






 Your current stack




 ChatGPT Plus
 $20/mo




 Claude Pro
 $20/mo




 Perplexity Pro
 $20/mo




 Gemini Advanced
 $20/mo




 X Premium+
 $16/mo






 Total / month
 $96








Suprmind Frontier



All five models · one thread · shared context





$95














Investment Analysts



### IC memo, defensible by 4pm



Five knowledge bases reference the same question. Build the strongest case for and against before capital gets committed.






 Research Symphony – pipeline




 01
 Retrieval

 47 sources cited





 02
 Analysis

 8 themes extracted





 03
 Fact-check

 3 contradictions flagged





 04
 Challenge

 Red-team pass





 05
 Synthesis

 8,200 / ~10,000 words












Companies solving day-to-day challenges with Suprmind




Real Work



## Built for people who need decisions that survive scrutiny.










> “5 AIs were a go-to resource in setting up our new business venture in NYC. From red teaming the initial idea (with harsh feedback), studio market and competitors analysis, to day to day brainstorming about launch phases and website setup. Being able to bounce any idea off 5 AIs, get a clear filtered answer and a todo list in 10 minutes helps a lot.”*LF




Luka Funduk



CEO, OFF Studio NYC & Funduck Production*> “I started using it for [competitor](https://suprmind.ai/hub/comparison/ai-fiesta-alternative/) research and it just kept expanding – new markets, risk reviews, compliance docs. Five different angles on the same question catches things I would have missed.”*AW




Aaron Weller



CEO & Co-founder, Miss Amara*> “We run everything through Suprmind now – new business ideas, client contracts, marketing strategies. Having five AIs push back on each other in one thread replaced hours of second-guessing between tools.”*MD




Milica D.



Co-founder & COO, Global Digital Marketing Agency*> “For analyzing business plans and evaluating client processes, the depth you get from five models reading each other is genuinely different. The Master Document export with custom prompt alone saves me hours on final reports.”*MT




Milos Tanasijevic



Senior International Adviser, EBRD – European Bank for Reconstruction and Development*5


Frontier Models






6


Orchestration Modes






25+


Master Document Templates






10K+


Words per Research Symphony Report









Disagreement is the feature.







## Radomir’s ideas get no special treatment. Yours won’t either.



Bring the plan you are most sure about. Five frontier AIs will read it, argue about it and tell you what breaks, before your money, your client or your board finds out.



 [Start Your Free Trial →](https://suprmind.ai/signup/spark)
 [Compare Plans](https://suprmind.ai/hub/pricing/)




7 days free. No credit card required. The trial runs GPT, Claude, Gemini and Grok. All five AIs, Debate and Red Team come with Pro at $45/mo.





FAQ



## Questions About the Lab







### What is Suprmind Data Lab?

 +





Suprmind Data Lab is the research arm of Suprmind. It publishes original data on how AI models behave: the Multi-Model AI Divergence Index, the AI Hallucination Rates & Benchmarks report and the AI Models Index. Every study is planned and argued out inside Suprmind, where five frontier AIs challenge the idea before any data work starts.









### Who does the research, people or AI?

 +





Both. Human researchers choose the questions, set the methods and review every report before it ships. AI does most of the legwork: five frontier models debate each study in Suprmind, and outside tools such as deep research agents gather and draft. When an AI and a primary source disagree, the source wins.









### Can I use your AI research charts in my article?

 +





Yes, with credit. AI Models Index charts are licensed CC BY 4.0, commercial use included, and come with embed code that updates with each new edition. The Divergence Index publishes its aggregate data under the same license. Credit Suprmind Data Lab and link to the report you quote.









### Is the underlying data available?

 +





Partly. The Multi-Model AI Divergence Index publishes 12 aggregate CSVs under CC BY 4.0, while turn-level data stays private to protect users. The AI Models Index dataset is not published, but you can request a custom cut at [[email protected]](/cdn-cgi/l/email-protection).









### How often is Data Lab research updated?

 +





It depends on the report. The AI Models Index ships a new edition every two weeks. The hallucination benchmarks report is revised as major benchmarks publish new results. The Divergence Index is an edition-based study built from Suprmind production data, and each edition carries its sample window and date.









### Does Suprmind’s own product bias the findings?

 +





It is a fair question, and the reports disclose it. Suprmind runs models from several of the labs we measure, so every report publishes its method, keeps weak numbers for providers in our own roster and names what the study cannot claim. The Divergence Index, for example, measures divergence, not ground truth.









### Where do the AI statistics on this page come from?

 +





Every research statistic on this page comes from one of the three Data Lab reports linked above, and each report states its sample, window and date. Third-party figures, such as LMArena ratings or AA-Omniscience scores, are credited to their original source inside the report that uses them.









### Can I run research like this in Suprmind myself?

 +





Yes. The lab works in the same [multi-AI platform](https://suprmind.ai/hub/platform/) you can try. Spark includes Sequential and Super Mind with a 7-day free trial and no credit card. Pro adds Debate, Red Team and First Principles, the modes the lab uses to trim ideas, and Research Symphony runs on Enterprise.








Disagreement is the feature.



Original AI research, argued over by five AIs first.

---

*Source: [https://suprmind.ai/hub/data-lab/](https://suprmind.ai/hub/data-lab/)*
*Generated by FAII AI Tracker v3.4.1*