Home Hub How It Works Features Use Cases How-To Guides Help Docs Pricing Login
Original AI Research, Made in a Five-AI Thread

Suprmind Data Lab:
Original AI Research
That Survived Five
Harsh but Fair AIs

Every Data Lab study starts as an idea pasted into a Suprmind thread, where five frontier AIs try to take it apart. The few that survive get scoped to real data, briefed, and run through Research Symphony. Three made it out so far: 1,324 production turns scored, 178 model releases tracked and six hallucination benchmarks cross-referenced.

  • Grok
  • Perplexity
  • Claude
  • ChatGPT
  • Gemini
Simplified · Sequential mode 5 models active
ChatGPT leans yes
Surface read says yes. The target could double our addressable market, and $42M equals 6x its stated ARR. On the headline numbers, the strategic case looks plausible.
Claude flag
The retention figure breaks the headline case. At 78% net revenue retention, the existing customer base contracts every year, so a 6x ARR price assumes growth the current economics do not support.
Perplexity correction
The precedent cited in the target's deck cannot be verified in the supplied files, so it should not anchor the valuation. The price case now depends on retention recovery and credible cross-sell evidence.
Gemini revised
Revising my read. Once the unverified precedent is removed and retention is priced correctly, the bull case loses its anchor. At $42M, the acquisition does not clear a reasonable diligence threshold.
Grok pushes back
Pushing back on the pile-on: weak retention does not make the asset worthless. A lower entry price, founder-retention earn-out and evidence of recoverable churn could preserve the strategic case. But those conditions need contractual proof.
ChatGPT re-priced
Updating the model. Pricing the retention risk and integration cost puts risk-adjusted fair value between $24M and $28M, well below $42M. Inside that band, the strategic upside begins to compensate for the downside.
Claude conditions
Two gates before close, whatever the price: NRR above 95% for two consecutive quarters, and an earn-out tied to team retention rather than headline revenue. Without both, a discount only buys the same problem cheaper.
Perplexity evidence
Using the scenario assumptions, a 3x to 4x ARR range is more defensible until retention recovers. On $7M ARR, that supports an offer between $21M and $28M rather than the requested $42M.
Gemini consensus
Converging: an entry near $26M sits inside the supported range. With the retention covenants and earn-out written into the deal, it survives the main objections raised in this thread.
Grok risk
One risk remains unpriced: the cost of walking. A rival acquisition could close the strategic window, so the board should compare that downside with the value destroyed by overpaying.
Master Document - Verdict
Do not acquire at $42M. Re-engage near $26M, conditional on two quarters of NRR recovery and a retention earn-out in writing. Otherwise, walk.
Type @ to mention one AI…

7 days free. No credit card. The trial runs GPT, Claude, Gemini and Grok.
All five AIs, Debate and Red Team come with Pro at $45/mo.

A small team of human researchers.
A much larger team of AIs.

The lab is a symbiotic team. Human researchers pick the questions, set the methods and sign off on what gets published. AI researchers do most of the heavy lifting: five frontier models arguing inside Suprmind, plus whatever outside AI tools a study needs, from the deep research agents each lab ships to analysis runs over leaderboard datasets.

Most of the thinking happens in Suprmind itself. Which question is worth asking, which data can answer it, what gets added and how to add it. All of it is argued out in shared threads, where every model reads what the others said before it answers.

The monthly bill for all those AI subscriptions, API keys and research platforms is responsible for most of Radomir’s sleepless nights. Radomir Basta, our founder and CEO, insists it is worth it. He insists quietly, usually around 2 a.m.

Human researchers
Set the question
Pick what is worth measuring, define the method, and check claims against primary sources before anything ships.
Five frontier AIs
Argue it out
GPT, Claude, Gemini, Grok and Perplexity in one Suprmind thread, challenging the premise, the data plan and the conclusions.
Outside AI tools
Do the legwork
Deep research agents, data tools and analysis runs, used wherever a study needs them. One study alone used five deep research AIs.
Primary sources
Settle the facts
Lab announcements, API changelogs, leaderboard datasets and benchmark papers. When an AI and a source disagree, the source wins.

Most research ideas die in the first thread.
That is the process working.

It usually starts with Radomir. He arrives with an idea that is bold, flashy and, by his own estimate, unlike anything the industry has seen. He pastes it into a Suprmind thread and asks five frontier AIs what they think.

They tell him. One finds the study that already covered it. One points out that the data to prove it does not exist. One asks what exactly would be measured, and how. By the end of the thread the idea is either smaller and sharper, or gone. Most are gone. That is why there are three Data Lab reports and not thirty.

The ego trimming is on purpose.

An idea that survives five harsh but fair critics is worth weeks of data work. An idea that does not was going to fail anyway, only later and at a higher price.

It is the same reason Suprmind exists. AI that agrees with you is pleasant. AI that tells you why you are wrong is useful.

  • Ideas argued out in shared Suprmind threads before any data work
  • Red Team and Debate modes aimed at the premise, not the polish
  • Data segments chosen to answer the question, not to decorate it
  • Research plans and briefs generated from the thread itself
  • Research Symphony for retrieval, analysis, fact-check, challenge and synthesis
  • Humans review every report and check claims against primary sources
1 The Idea Arrives Usually Radomir’s
Bold, flashy and often late at night. Pasted into a Suprmind thread before anyone can talk him out of it.
2 Five AIs Push Back Ego Trimming
Red Team and Debate test novelty, feasibility and data. Most ideas end here, politely.
3 The Survivor Gets Focused Scope and Data
What is left gets narrowed to specific data segments that can actually answer the question.
4 Plan and Briefs Written From the Thread
The conversation becomes a research plan and briefs, built from what the models settled.
5 Research Symphony Runs Then Humans Check
Retrieval, analysis, fact-check, challenge and synthesis. Then humans check the claims against primary sources.

AI statistics we measured ourselves.
Not copied from a press release.

Three studies, three questions: how often frontier AIs contradict each other in real work, how often they make things up, and whether new models actually beat the ones they replace. Every report publishes its method, and every number is there to cite.

001

Original research

Multi-Model AI Divergence Index

April 2026 Edition – The Confidence Trap

Suprmind’s own production data. 1,324 multi-AI turns from 299 users over 45 days, scored for contradiction, correction and unique insight per provider. Where five frontier AIs disagree, who catches whom, and how often a confident answer fails peer review.

51.3%

Gemini’s confident answers contradicted or corrected

9.77×

Perplexity vs Gemini catch ratio

72.1%

Disagreement on financial questions

Published: April 2026 Sample: 1,324 production turns Users: 299 Window: 45 days License: CC BY 4.0, 12 CSVs
Read the research
002

Live benchmark

AI Hallucination Rates & Benchmarks

Living reference, revised as benchmarks publish

Seven hallucination benchmarks in one report: Vectara HHEM, AA-Omniscience, FACTS, HalluHard, CJR Citation, SimpleQA and SimpleQA Verified. Cross-referenced model by model and read alongside Suprmind’s own multi-model findings.

61 / 202

Models with a positive AA-Omniscience Index

72.6%

Claude Fable 5.1 wrong instead of refusing

73-86%

Drop with web search on, in OpenAI’s own evals

Benchmarks: 7 Models: 202 on AA-Omniscience Sources: 50+ Latest revision: October 2026 Format: Open access
Read the research
003

Live index

AI Models Index

October 2026 Edition, updated every two weeks

Every premium AI model from 15 labs since 2023, with the day the public could first use it and a blind-vote check on whether it beat the model it replaced. 178 releases, 78 version pairs measured on LMArena and nine live charts you can embed.

4.5 days

Between premium model releases in 2026

56%

Of measured 2026 releases beat the model they replaced

1 in 6

Measured 2026 releases that lost to their predecessor

Releases: 178 Labs: 15 Refresh: Every two weeks Charts: CC BY 4.0 Embeds: 9 live charts
Read the research

A confident AI answer
is not the same as a correct one.

The Multi-Model AI Divergence Index measures what happens when several frontier models answer the same question in the same thread. Over 45 days, 1,324 production turns from 299 users, with two to five models per turn, ran through Suprmind’s Disagreement/Correction Index classifier, which scored each provider for confidence, contradiction, correction and unique insight.

The headline finding named the report. Confidence and accuracy are different signals. Gemini’s high-confidence answers were contradicted or corrected by another model 51.3% of the time, the highest of the five. When the stakes rose, Claude’s contradicted-while-confident rate fell 7.5 points. Gemini’s barely moved.

Bar chart of AI catch ratios per provider: Perplexity 2.54, Claude 2.25, Grok 0.72, GPT 0.38, Gemini 0.26
Above 1.0, a provider catches more errors than it gets caught making. Perplexity and Claude made 60.7% of the corrections attributable to a specific model.
Peer corrections
1,401
Logged across 1,324 turns, about one per turn. Each one a claim another model refused to let stand.
Unique insights
3,484
Points one model raised that no other model in the thread did. Five models, five different blind spots.
Turns with signal
99.1%
Turns that surfaced at least one contradiction, correction or unique insight. Only 0.9% stayed silent.
Money beats law
72.1%
Disagreement on financial questions, against 41.2% on legal ones. The models fight 1.75× more over money.

What it does not measure: ground truth. The classifier tracks where models diverge, not which one is right. Divergence tells you where to look. Evidence tells you who was right.

Six hallucination benchmarks.
One report. No spin.

Every major AI model hallucinates, and every benchmark measures it differently. Vectara tests summarization. AA-Omniscience tests whether a model answers when it should refuse. FACTS, HalluHard, CJR Citation, SimpleQA and SimpleQA Verified each catch a different failure. Read one leaderboard and you see one slice.

The AI Hallucination Rates & Benchmarks report puts them side by side: 202 models on the AA-Omniscience board, frontier models profiled one by one, 50+ cited sources and a cross-benchmark table revised as new results publish. The numbers disagree with each other, and the report explains why that matters more than any single ranking.

Screenshot of the Suprmind cross-benchmark AI hallucination rates table, August 2026 revision
The cross-benchmark reference table, August 2026 revision. The live table covers every frontier model in the report.
Reliability
61 / 202
Models whose correct answers outnumber their wrong ones on AA-Omniscience. The other 141 guess more than they know.
Web search effect
73-86%
Drop in hallucination when web search is switched on, across OpenAI’s own factuality evaluations.
Lowest rate
0.6%
Best score on Vectara’s original summarization dataset. Low on one test does not mean low on all seven.
One update
38 pts
How far Gemini 3.1 Pro cut its AA-Omniscience hallucination rate in a single release, from 88% to 50%.

New AI models ship faster.
They improve less.

The AI Models Index tracks every premium model released by 15 leading labs since 2023. Each release gets two answers: the first day the public could actually use it, and whether people preferred it over the version it replaced in blind head-to-head votes on LMArena.

It is also the study where the lab put its own tools on trial. The same research brief went to five deep research AIs, and each built a release history on its own. None got it right alone. Two called a 2026 slowdown that never happened. Every disagreement was settled against lab announcements, API changelogs and LMArena’s published data.

Live chart from the AI Models Index. It updates with every edition, and the Embed button on the index gives you the code for your own site.
Release pace
4.5 days
Between premium model releases across the industry in 2026, down from 7.5 days in 2023.
Improvement rate
56%
Of measured 2026 releases beat the model they replaced in blind votes. In 2025 it was 74%.
Regression rate
1 in 6
Measured 2026 releases that lost to the model they replaced by a statistically real margin.
Smaller steps
+9.3
Median LMArena gain per 2026 release, down from +21.5 in 2025.

“A premium AI model now ships every four and a half days, but this year only a little over half of them beat the version they replaced in blind votes, and one in six lost. The newest model is a candidate, not an upgrade. Test it on your own work before you switch.”

Radomir Basta, founder and CEO of Suprmind

Rules the lab follows,
even when the data is awkward for us.

Methods in the open

Every report publishes its definitions and method in full, so you can check how each number was made.

Bad numbers stay in

Several labs we measure supply the models Suprmind runs. Their weak results get published the same as their strong ones.

Limits named up front

The Divergence Index does not measure ground truth. LMArena preference is not task performance. Each report says what it cannot claim.

Sources settle it

AI tools gather and draft. Lab announcements, changelogs and benchmark datasets decide what is true.

Disclosures on the page

When Claude scored five deep research runs, including its own, the report said so. The scoring rules were fixed before any scoring started.

Open by default

AI Models Index charts are licensed CC BY 4.0. The Divergence Index ships 12 aggregate CSVs under the same license. Turn-level data stays private.

Cite it, embed it,
or ask us for a custom cut.

Data Lab research is written to be reused. AI Models Index charts and the Divergence Index dataset are licensed CC BY 4.0, so you can use them, including commercially, with credit to Suprmind Data Lab and a link to the source report. For interviews, press kits and data requests, write to [email protected].

Cite

Suprmind Data Lab, report name and edition, with a link to the report page. Credit LMArena and Artificial Analysis for their own figures.

Embed live charts

Every AI Models Index chart has an Embed button. Embedded charts update with each edition, so your page never quotes stale numbers.

Download data

The Divergence Index publishes 12 aggregate CSVs under CC BY 4.0. The AI Models Index dataset is not published, but custom cuts are available on request.

Press kit

Fact sheet, light and dark charts, the race-for-#1 video and logos. Request it through the press kit form at the end of the AI Models Index.

Every study runs on the same
six modes Suprmind ships.

Red Team and Debate do the ego trimming. Sequential builds the analysis. First Principles rebuilds the question when the obvious framing fails. Research Symphony runs the full pipeline. Switch modes mid-thread and every model keeps the context.

Sequential

Default

AIs respond one after another. Each reads everything before it. The default and the deepest.

Best for:

Complex analysis, research, architecture decisions

Learn more
You Doc

Super Mind

Fastest

All five respond simultaneously. A sixth AI synthesizes one unified answer with consensus and divergence mapped.

Best for:

Quick decisions, fact verification, time-sensitive calls

Learn more
You Doc

Debate

AIs argue assigned positions in sequence. Rebuttals and counter-arguments. Minority views preserved.

Best for:

Strategy validation, thesis stress-testing

Learn more
You ×3 Doc

Red Team

AIs attack your plan from six angles in sequence: financial, technical, reputational, regulatory, operational, edge cases.

Best for:

Pre-launch validation, risk assessment, investment pre-mortems

Learn more
You Doc

Research Symphony

Enterprise

Automated research pipeline that retrieves sources, analyses, fact-checks, challenges, and synthesises. Produces 10,000+ word reports with citations.

Best for:

Deep research, comprehensive reports

Learn more
You Doc

First Principles

Pro+

Strips a question to its fundamentals. Each model names its assumptions, identifies the underlying axioms, then rebuilds the analysis from the ground up.

Best for:

Highest-stakes decisions where convention is suspect

You Doc

Sequential, Debate, Red Team, and First Principles all use sequential orchestration – each AI builds on what came before. Super Mind mode runs in parallel with a synthesis layer. Chain any combination mid-conversation.

Beyond the Lab

Four jobs, four shipped artifacts.

Research is one job the five AIs do. Here are four more, and every output is a real document you can export and send.

Strategy Consultants

M&A pre-mortem in 90 minutes

Walk into the partner meeting with five frontier AIs already disagreeing on your behalf. Each fabrication caught before slides leave your laptop.

Master Document – preview v4 · exported as PDF

Skybridge Acquisition – Recommendation Memo

Prepared by Suprmind · Sequential mode · 5 models · 47 min

Verdict

Do not acquire at $42M. Revisit at $26M with NRR turnaround proof.

Executive summary
Five-model consensus matrix
Disagreements & unresolved questions
Risk register (red team output)
Supporting evidence – citations

Founders & Operators

Pricing experiment, defended

Run a $79 vs $149 split through Debate mode. Watch Claude argue retention, Grok argue elasticity, Perplexity ground both in 2026 benchmarks.

Debate transcript – preview
Claude PRO – $149

Retention curve flattens past $99. The $50 of headroom buys you Frontier-buyer signaling.

Grok CON – $79

Elasticity at this stage is brutal. You’ll lose 31% of conversions for ~22% revenue lift.

Perplexity CONTEXT

2026 SaaS prosumer benchmarks: 38% of $99+ tools see >40% trial-to-paid lift after price reduction.

AI Power Users

Stop reconciling five tabs

Cancel ChatGPT Pro, Claude Pro, Perplexity Pro, Gemini Advanced. One conversation. Five models. Shared context. $95/mo all-in.

Your current stack
ChatGPT Plus $20/mo
Claude Pro $20/mo
Perplexity Pro $20/mo
Gemini Advanced $20/mo
X Premium+ $16/mo
Total / month $96

Suprmind Frontier

All five models · one thread · shared context

$95

Investment Analysts

IC memo, defensible by 4pm

Five knowledge bases reference the same question. Build the strongest case for and against before capital gets committed.

Research Symphony – pipeline
01 Retrieval 47 sources cited
02 Analysis 8 themes extracted
03 Fact-check 3 contradictions flagged
04 Challenge Red-team pass
05 Synthesis 8,200 / ~10,000 words

Built for people who need decisions
that survive scrutiny.

“I started using it for competitor research and it just kept expanding – new markets, risk reviews, compliance docs. Five different angles on the same question catches things I would have missed.”

Aaron Weller

CEO & Co-founder, Miss Amara

“We run everything through Suprmind now – new business ideas, client contracts, marketing strategies. Having five AIs push back on each other in one thread replaced hours of second-guessing between tools.”

Milica D.

Co-founder & COO, Global Digital Marketing Agency

“For analyzing business plans and evaluating client processes, the depth you get from five models reading each other is genuinely different. The Master Document export with custom prompt alone saves me hours on final reports.”

Milos Tanasijevic

Senior International Adviser, EBRD – European Bank for Reconstruction and Development

5
Frontier Models
6
Orchestration Modes
25+
Master Document Templates
10K+
Words per Research Symphony Report

Disagreement is the feature.

Radomir’s ideas get no special treatment.
Yours won’t either.

Bring the plan you are most sure about. Five frontier AIs will read it, argue about it and tell you what breaks, before your money, your client or your board finds out.

7 days free. No credit card required. The trial runs GPT, Claude, Gemini and Grok. All five AIs, Debate and Red Team come with Pro at $45/mo.

Questions About the Lab

What is Suprmind Data Lab?

Suprmind Data Lab is the research arm of Suprmind. It publishes original data on how AI models behave: the Multi-Model AI Divergence Index, the AI Hallucination Rates & Benchmarks report and the AI Models Index. Every study is planned and argued out inside Suprmind, where five frontier AIs challenge the idea before any data work starts.

Who does the research, people or AI?

Both. Human researchers choose the questions, set the methods and review every report before it ships. AI does most of the legwork: five frontier models debate each study in Suprmind, and outside tools such as deep research agents gather and draft. When an AI and a primary source disagree, the source wins.

Can I use your AI research charts in my article?

Yes, with credit. AI Models Index charts are licensed CC BY 4.0, commercial use included, and come with embed code that updates with each new edition. The Divergence Index publishes its aggregate data under the same license. Credit Suprmind Data Lab and link to the report you quote.

Is the underlying data available?

Partly. The Multi-Model AI Divergence Index publishes 12 aggregate CSVs under CC BY 4.0, while turn-level data stays private to protect users. The AI Models Index dataset is not published, but you can request a custom cut at [email protected].

How often is Data Lab research updated?

It depends on the report. The AI Models Index ships a new edition every two weeks. The hallucination benchmarks report is revised as major benchmarks publish new results. The Divergence Index is an edition-based study built from Suprmind production data, and each edition carries its sample window and date.

Does Suprmind’s own product bias the findings?

It is a fair question, and the reports disclose it. Suprmind runs models from several of the labs we measure, so every report publishes its method, keeps weak numbers for providers in our own roster and names what the study cannot claim. The Divergence Index, for example, measures divergence, not ground truth.

Where do the AI statistics on this page come from?

Every research statistic on this page comes from one of the three Data Lab reports linked above, and each report states its sample, window and date. Third-party figures, such as LMArena ratings or AA-Omniscience scores, are credited to their original source inside the report that uses them.

Can I run research like this in Suprmind myself?

Yes. The lab works in the same multi-AI platform you can try. Spark includes Sequential and Super Mind with a 7-day free trial and no credit card. Pro adds Debate, Red Team and First Principles, the modes the lab uses to trim ideas, and Research Symphony runs on Enterprise.

Disagreement is the feature.

Original AI research, argued over by five AIs first.