Suprmind Data Lab · Last updated on October 4, 2026
Every premium AI model from 15 labs since 2023, with its verified release date and a blind-vote check on whether it beat the model it replaced.
Updated every two weeks.
October 4, 2026 edition: GPT-6.1 Sol, Claude Sonnet 5.5, Claude Opus 5.5, GPT-6 Sol and Grok 4.7 added. Gemini 4 Argon took #1 on LMArena on September 30 and is still limited to trusted testers. 178 releases tracked, 78 version pairs measured.
LMArena has published 187 leaderboard snapshots since August 2024. The top spot changed hands 3 times in late 2024, 9 times in 2025 and 12 times in 2026 so far. Press play to watch the top 10 reshuffle, or drag the slider to any date.
OpenAI held #1 the longest in total, then Anthropic, then Google. Claude Fable 5 had the longest single reign of 2026 at 107 days. Right now the #1 model is Gemini 4 Argon, which took the top spot on September 30 and is still restricted to trusted testers. Below it, the top three public models sit within one rating point of each other, a statistical tie.
Frontier labs now ship a premium AI model somewhere every 4.5 days. Version numbers climb faster than the models improve. In 2026, 33 of 53 premium releases were point releases, and the typical release beat its predecessor by less than half the margin it managed a year earlier.
Suprmind Data Lab built this index the way Suprmind works. Five deep research AIs - Claude, ChatGPT, Gemini, Grok and Perplexity - researched the release history independently. Every disagreement between them was then checked against primary sources: lab announcements, API changelogs and LMArena's published leaderboard data. None of the five got it right alone, and you can see how each one did.
Each release gets two answers. When did it really become available to the public? And did people actually prefer it over the version before? The second answer comes from blind head-to-head votes, not from launch posts.
2026. A premium release every 4.5 days, and labs logged 51 release days by October 3, against 42 at the same point in 2025.
Anthropic - a median of 22.5 days between premium releases in 2026, down from 133 days in 2023.
OpenAI - a median of 52 to 57 days between premium releases every year since 2025.
Google and DeepSeek. Google's median gap grew from 43.5 days in 2025 to 79.5 in 2026, DeepSeek's from 63 to 127.5.
Mistral Medium 3 over Mistral Medium, +164.9 LMArena points after 513 days. It wins 72.1% of head-to-head votes.
Grok 4.3 against Grok 4.20, -33.2 points after 59 days. GPT-6 Sol is second, -26.7 against GPT-5.6 Sol.
Point releases and snapshots shipped within 45 days gained a median of +0.1 points. Those after longer gaps gained +14.9.
People prefer GPT-5.1. GPT-5.2 wins 47.3% of blind votes, a statistically real loss.
Gemini 4 Argon, which the public cannot use yet. The top three public models sit within one rating point of each other.
Claude Fable 5 held #1 for 107 days in 2026. In total since August 2024: OpenAI 314 days, Anthropic 240, Google 212.
GPT-5.2's launch post led with +35.3 points on ARC-AGI-2. People rated it 18.5 LMArena points below GPT-5.1.
ChatGPT, with 94% of release dates exactly right. Gemini's LMArena figures matched 1 of 12.
Faster every year. Counting premium releases from 15 labs that set or chase the frontier, the industry shipped 15 in 2023, 50 in 2024, 60 in 2025 and 53 in 2026 through October 3. Measured over the same window each year, January 1 to October 3, and counting each lab's release days separately, release days went from 6 to 37 to 42 to 51. That puts 2026 21% ahead of 2025 at the same point.
| Year | Releases | One every | Median lab gap | New generations | Point releases |
|---|---|---|---|---|---|
| 2023 | 15 | 7.5 days | 133 days | 12 of 15 | 1 |
| 2024 | 50 | 6.5 days | 84.5 days | 18 of 50 | 10 |
| 2025 | 60 | 6 days | 83 days | 19 of 60 | 16 |
| 2026 (to Oct 3) | 53 | 4.5 days | 59 days | 10 of 53 | 33 |
Version numbers are outrunning new generations. In 2023, 12 of 15 premium releases were new generations. In 2026, 33 of 53 were point releases - a 5.1 becoming a 5.2 - and only 10 were new generations.
"A new model every few weeks" is true for the industry and false for most labs. A typical leading lab ships every six to twelve weeks. The pace comes from the number of labs shipping at once, plus one outlier covered in section 4.
Two lines tell the story. Release volume keeps climbing. The gain people feel from each release keeps shrinking. In 2025 the median release beat its predecessor by +21.5 LMArena points, after a median wait of 120 days. In 2026 it wins by +9.3, after 75 days. Among Western labs the 2026 median is +2.5, and 10 of 29 measured releases went backwards.
Top: premium releases from 15 labs over a trailing six months. Bottom: median LMArena rating change against the predecessor over the same window, text leaderboard with style control. Hover for monthly values.
Two cautions keep this honest. LMArena compresses at the top of its scale: when every model is strong, preference gaps shrink, so part of the smaller step is a measurement effect. And capability still rises. On the Artificial Analysis Intelligence Index, none of the 39 version pairs we could compare went backwards. Point releases gained a median of 4 index points and new generations 5.5.
First impressions fade too. In 45 of 76 version pairs with more than one LMArena snapshot, the newer model's lead is smaller today than in its launch week. GPT-5 led GPT-4.5 by 43.3 points at launch. Today it trails by 10.2.
Anthropic, by a wide margin. Its median gap between premium releases fell from 133 days in 2023 to 22.5 days in 2026, across ten releases in its Opus, Sonnet and Fable lines. Every other leading lab sits between six weeks and five months.
| Lab | 2023 | 2024 | 2025 | 2026 | 2026 releases |
|---|---|---|---|---|---|
| Anthropic | 133 | 108 | 75 | 22.5 | 10 |
| Alibaba (Qwen) | - | 75 | 83 | 45.5 | 6 |
| MiniMax | - | - | 142.5 | 51 | 3 |
| OpenAI | 118.5 | 69 | 56.5 | 52.5 | 8 |
| Zhipu AI (Z.ai) | - | 226 | 83 | 56.5 | 4 |
| xAI | - | 121 | 131 | 59 | 5 |
| Meta | - | 136 | 120 | 60 | 4 |
| Google DeepMind | 210 | 58 | 43.5 | 79.5 | 4 |
| Moonshot AI | - | - | 117 | 83 | 3 |
| DeepSeek | - | 55.5 | 63 | 127.5 | 2 |
| Mistral AI | 75 | 117 | 112 | 147 | 1 |
Median days between a lab's premium releases, counted in release days, so same-day siblings like GPT-5.4 and GPT-5.4 Pro do not create zero-day gaps.
OpenAI is the steady one, between 52 and 57 days every year since 2025. Google and DeepSeek slowed down in 2026. Google's Pro line has sat at Gemini 3.1 Pro since February, a teased Gemini 3.5 Pro never reached the API changelog, and Gemini 4 Argon is still limited to trusted testers. DeepSeek shipped twice in 2026, 111 days apart.
Every premium release from 15 labs, placed on its first day of public use in an app or API. The color on each model is the verdict against the model it replaced, and the number is its LMArena score change. Hover or tap a model for the predecessor, the win rate and the source. The impact chart under it puts every step since 2023 on one axis.
No premium release in over a year: Microsoft (last Apr 2025), Amazon (last Apr 2025).
Scroll sideways for more months. Tap any model for details.
Gray models were never rated against their predecessor on LMArena, mostly 2023 and 2024 releases and gated models. Next release windows are estimates: each lab's last premium release plus the median of its last five gaps. They are not announced dates.
The same releases on one time axis. Bigger dots won more blind votes against the model they replaced, and color shows the verdict. Click a verdict to hide it, or narrow the view to 2025 or 2026.
Hollow dots are releases where one of the two models never appeared on LMArena, mostly 2023 and 2024 models and gated releases. LMArena's public leaderboard data starts in August 2024.
The full ledger runs from the first premium releases of 2023 to the latest AI models in this edition. Every chip carries the lab, release date, access route, predecessor and source. Hover or tap one to see them.
Pick the model you use now and the one you are considering. The comparison uses blind head-to-head votes from LMArena, not launch claims. For direct successors you also get verified notes on what changed in daily use - speed, price, token use, hallucination and behavior - labelled by evidence type.
Read the verdict for what it is. LMArena measures which answer people prefer in open-ended chat. A model can lose there and still be the better pick for agentic coding or long tool chains, which is why the Intelligence Index change sits next to it wherever a like-for-like comparison exists.
Application
A new premium model lands every few days, and the newest is not always the right one for your task. Suprmind puts five frontier AIs in one conversation, and new models reach its AI Teams within days of release. When one model regresses on your question, the others read its answer and say so before you act on it.
See how the multi-AI workflow worksPoint releases and snapshots that ship within 45 days of the previous version gained a median of +0.1 points. 5 of the 11 lost, and every one of those losses is statistically real. GPT-5.2 arrived 29 days after GPT-5.1 and wins 47.3% of votes against it. Claude Opus 4.8, Grok 4.6, Grok 4.7 and Qwen3-Max show the same pattern.
Fast is not always bad. GPT-6.1 Sol shipped 7 days after GPT-6 Sol and gained 26 points, mostly by fixing what GPT-6 Sol lost against its own predecessor. MiniMax M2.7 gained 24.1 after 34 days.
The time since the last release explains more than the version label. Compare releases that arrived after the same wait, and the label stops mattering:
| Gap since previous version | New generation | Point release | Snapshot |
|---|---|---|---|
| 46 to 90 days | +13.2 (7 pairs) | +10.8 (19 pairs) | +10.2 (4 pairs) |
Median LMArena change against the predecessor. Generations look bigger overall mainly because they arrive after longer waits.
Launch posts lead with the benchmark that moved most. Across 19 version pairs measured both ways, the median of the biggest gain each lab chose to report was +16.5 points. The same pairs moved +3 on the Artificial Analysis Intelligence Index. The two numbers sit on different scales, so read this as which number the launch post picked, not as a ratio.
None of this makes the launch numbers false. A model can get far better at competition math or agentic coding and feel no better, or worse, in conversation. The practical point: test a new version on your own work before you switch.
Versions ship faster than anyone can re-test them. In Suprmind, five frontier AIs answer in the same thread and read each other's work, so when the newest model gets your question wrong, another one catches it before you act on it.
Test It on a Real Question7 days free. No credit card. Grok, GPT, Claude and Gemini in the trial.
Monday to Thursday, almost always. Of 113 premium releases since January 2025, Thursday leads with 28, Friday has 8 and weekends 5. The busiest week on record was August 31 to September 6, 2026: Claude Fable 5.1, Gemini 3.8 Flash, Muse Spark 1.3, GPT-6 Astra and a Qwen3.8-Max snapshot shipped within three days.
Premium releases by weekday, January 2025 to October 3, 2026.
The theory that labs time launches to counter each other does not hold up. 103 of the 113 premium releases since January 2025, or 91%, landed within seven days of another lab's release. Spread the same releases over random dates and the share is still 91%.
Announced is not shipped. Google has the longest waits between announcing a model and letting the public use it: 73 days for Gemini 2.5 Deep Think, 64 for Gemini 1.0 Ultra and 54 for Gemini 1.5 Pro. This index dates every model by first public use, which is why Gemini 4 Argon is not in the ledger yet.
We gave the same research brief to five deep research AIs and scored every answer against primary sources. The results are the Confidence Trap in fresh form. Two platforms called a 2026 slowdown that never happened, because each missed 22 or more of the year's 53 releases. One produced LMArena figures that match no snapshot LMArena has ever published. Only 49 of 195 models appeared in all five ledgers.
The AIs also know their own makers unevenly. ChatGPT and Claude got every release date for OpenAI and Anthropic exactly right. Grok got xAI's dates exactly right half the time, against 92% for other labs. Gemini found 11 of Google's 21 premium releases.
Disclosure. Claude scored this, including Claude's own run. The rules were fixed before scoring, and the scoring is mechanical.
Five frontier models in one thread, each reading the answers before it. Test it on a question where being wrong costs you.
Try Multi-AI ValidationLMArena preference is not task performance, and a statistical tie at the top is a tie. Dates for Alibaba, ByteDance and Moonshot rest partly on secondary sources. Silent updates behind the same model alias are not counted. Gemini experimental checkpoints never released to the public are not in the ledger.
You can quote any number or chart on this page with a link to the AI Models Index. Capability figures credit Artificial Analysis. Charts are free to use under CC BY 4.0. Every chart has Download, Embed and Cite buttons, and embedded charts update themselves with each edition.
Cite it as: Suprmind Data Lab, AI Models Index, October 4, 2026 edition, with a link to this page.
High-resolution charts in light and dark, a fact sheet with the key numbers and how each one is measured, the method on one page, logos and a founder quote. The link appears as soon as you submit, and a copy goes to your inbox.
Your press kit is ready. We also sent the link to your inbox.
Open the press kitQuestions or a custom cut of the data: [email protected]. We handle your details under our privacy policy.
In 2026 a premium AI model ships somewhere every 4.5 days, counting 15 leading labs. That is up from every 7.5 days in 2023. Individual labs are slower: a typical leading lab ships a premium model every six to twelve weeks, and Anthropic is the fastest at a median of 22.5 days in 2026.
Anthropic. Its median gap between premium releases was 22.5 days in 2026, across ten releases in the Opus, Sonnet and Fable lines. Alibaba's Qwen team was next at 45.5 days, then MiniMax at 51 and OpenAI at 52.5. Google and DeepSeek slowed down in 2026.
Usually, but by less each year. The median new release beat its predecessor by 21.5 LMArena points in 2025 and by 9.3 in 2026. Of 43 measured releases in 2026, 10 scored below the version they replaced, 7 of them by a statistically real margin. Capability benchmarks still rise: none of 39 version pairs went backwards on the Artificial Analysis Intelligence Index.
Not for open-ended chat. In LMArena's blind head-to-head votes, GPT-5.2 wins 47.3% against GPT-5.1, an 18.5-point gap that is statistically real. OpenAI's launch post reported large gains on reasoning benchmarks such as ARC-AGI-2, so GPT-5.2 can still be the better pick for some tasks. Test it on your own work before you switch.
There is no single answer. On LMArena's text leaderboard of October 2, 2026, Gemini 4 Argon holds #1 but is limited to trusted testers. The top three public models, Claude Opus 4.6, Claude Fable 5 and Claude Opus 5.5, sit within one rating point of each other, a statistical tie, and no model leads every task.
Each release is compared with the model it replaced: usually the previous version in the same model line, otherwise the lab's previous flagship. The comparison uses LMArena's text leaderboard with style control and the highest-rated variant of each model. The rating gap converts into the share of blind votes the newer model would win. A loss counts when it clears the 95% confidence band, and a gain when it also wins at least 51% of votes. Where a like-for-like comparison exists, the Artificial Analysis Intelligence Index adds a capability view.
A model the lab positioned as its most capable at release, or a new version of a flagship line such as Claude Opus, Gemini Pro or GPT. The release date is the first day the public could use it in an app or API. Waitlists, trusted-tester programs and arena-only testing do not count.
A new edition ships every two weeks. Each edition adds new releases, refreshes the LMArena data, measures new version pairs and re-checks dates against lab announcements and API changelogs.
Yes, with attribution. Quote any number or chart with a link to the AI Models Index, and credit Artificial Analysis for Intelligence Index figures. Charts are free to use under CC BY 4.0. Every chart has Download, Embed and Cite buttons, and embedded charts update themselves with each edition. For the press kit with high-resolution charts and a fact sheet, use the press kit form or email [email protected].
Five frontier models. One conversation. Every answer read and challenged by the next model in the thread. When a release disappoints, you find out on your question, not in next month's leaderboard.
Compare Plans