Home Hub How It Works Features Use Cases How-To Guides Help Docs Pricing Login

Table of Contents

Suprmind Data Lab · Last updated on October 4, 2026

AI Models Index: Tracking New AI Model
Releases and Improvement Rates

Every premium AI model from 15 labs since 2023, with its verified release date and a blind-vote check on whether it beat the model it replaced.
Updated every two weeks.

October 4, 2026 edition: GPT-6.1 Sol, Claude Sonnet 5.5, Claude Opus 5.5, GPT-6 Sol and Grok 4.7 added. Gemini 4 Argon took #1 on LMArena on September 30 and is still limited to trusted testers. 178 releases tracked, 78 version pairs measured.

4.5 days
Between premium model releases across the industry in 2026. It was 7.5 days in 2023.
56%
Improvement rate in 2026: measured releases that clearly beat the model they replaced in blind votes. It was 74% in 2025.
16%
Regression rate in 2026: about 1 in 6 measured releases clearly lost to the model they replaced. It was 11% in 2025.
+9.3
Median LMArena gain per release in 2026, down from +21.5 in 2025

The Race for #1 on LMArena

LMArena has published 187 leaderboard snapshots since August 2024. The top spot changed hands 3 times in late 2024, 9 times in 2025 and 12 times in 2026 so far. Press play to watch the top 10 reshuffle, or drag the slider to any date.

Who held #1, and for how long

OpenAI held #1 the longest in total, then Anthropic, then Google. Claude Fable 5 had the longest single reign of 2026 at 107 days. Right now the #1 model is Gemini 4 Argon, which took the top spot on September 30 and is still restricted to trusted testers. Below it, the top three public models sit within one rating point of each other, a statistical tie.

From the research behind this edition
22 of 53
2026 releases missed by the two deep research AIs that wrongly called a 2026 slowdown
1 of 12
LMArena scores from Gemini's deep research that match any snapshot LMArena has published
45 of 76
Version pairs where the launch-week lead on LMArena has since shrunk
12
Times the #1 spot on LMArena changed hands in 2026, up from 9 in all of 2025

Frontier labs now ship a premium AI model somewhere every 4.5 days. Version numbers climb faster than the models improve. In 2026, 33 of 53 premium releases were point releases, and the typical release beat its predecessor by less than half the margin it managed a year earlier.

Suprmind Data Lab built this index the way Suprmind works. Five deep research AIs - Claude, ChatGPT, Gemini, Grok and Perplexity - researched the release history independently. Every disagreement between them was then checked against primary sources: lab announcements, API changelogs and LMArena's published leaderboard data. None of the five got it right alone, and you can see how each one did.

Each release gets two answers. When did it really become available to the public? And did people actually prefer it over the version before? The second answer comes from blind head-to-head votes, not from launch posts.

Quick-Reference Findings

Fastest year on record

2026. A premium release every 4.5 days, and labs logged 51 release days by October 3, against 42 at the same point in 2025.

Fastest-shipping lab

Anthropic - a median of 22.5 days between premium releases in 2026, down from 133 days in 2023.

Steadiest lab

OpenAI - a median of 52 to 57 days between premium releases every year since 2025.

Slowing down

Google and DeepSeek. Google's median gap grew from 43.5 days in 2025 to 79.5 in 2026, DeepSeek's from 63 to 127.5.

Biggest measured step

Mistral Medium 3 over Mistral Medium, +164.9 LMArena points after 513 days. It wins 72.1% of head-to-head votes.

Biggest regression

Grok 4.3 against Grok 4.20, -33.2 points after 59 days. GPT-6 Sol is second, -26.7 against GPT-5.6 Sol.

Fast follow-ups

Point releases and snapshots shipped within 45 days gained a median of +0.1 points. Those after longer gaps gained +14.9.

GPT-5.2 vs GPT-5.1

People prefer GPT-5.1. GPT-5.2 wins 47.3% of blind votes, a statistically real loss.

Current #1 on LMArena

Gemini 4 Argon, which the public cannot use yet. The top three public models sit within one rating point of each other.

Longest at #1

Claude Fable 5 held #1 for 107 days in 2026. In total since August 2024: OpenAI 314 days, Anthropic 240, Google 212.

Launch number vs blind votes

GPT-5.2's launch post led with +35.3 points on ARC-AGI-2. People rated it 18.5 LMArena points below GPT-5.1.

Best AI at researching AI

ChatGPT, with 94% of release dates exactly right. Gemini's LMArena figures matched 1 of 12.

How Often Do AI Labs Release New Models?

Faster every year. Counting premium releases from 15 labs that set or chase the frontier, the industry shipped 15 in 2023, 50 in 2024, 60 in 2025 and 53 in 2026 through October 3. Measured over the same window each year, January 1 to October 3, and counting each lab's release days separately, release days went from 6 to 37 to 42 to 51. That puts 2026 21% ahead of 2025 at the same point.

YearReleasesOne everyMedian lab gapNew generationsPoint releases
2023157.5 days133 days12 of 151
2024506.5 days84.5 days18 of 5010
2025606 days83 days19 of 6016
2026 (to Oct 3)534.5 days59 days10 of 5333

Version numbers are outrunning new generations. In 2023, 12 of 15 premium releases were new generations. In 2026, 33 of 53 were point releases - a 5.1 becoming a 5.2 - and only 10 were new generations.

"A new model every few weeks" is true for the industry and false for most labs. A typical leading lab ships every six to twelve weeks. The pace comes from the number of labs shipping at once, plus one outlier covered in section 4.

The Upgrade Treadmill: Faster Releases, Smaller Steps

Two lines tell the story. Release volume keeps climbing. The gain people feel from each release keeps shrinking. In 2025 the median release beat its predecessor by +21.5 LMArena points, after a median wait of 120 days. In 2026 it wins by +9.3, after 75 days. Among Western labs the 2026 median is +2.5, and 10 of 29 measured releases went backwards.

Top: premium releases from 15 labs over a trailing six months. Bottom: median LMArena rating change against the predecessor over the same window, text leaderboard with style control. Hover for monthly values.

Two cautions keep this honest. LMArena compresses at the top of its scale: when every model is strong, preference gaps shrink, so part of the smaller step is a measurement effect. And capability still rises. On the Artificial Analysis Intelligence Index, none of the 39 version pairs we could compare went backwards. Point releases gained a median of 4 index points and new generations 5.5.

First impressions fade too. In 45 of 76 version pairs with more than one LMArena snapshot, the newer model's lead is smaller today than in its launch week. GPT-5 led GPT-4.5 by 43.3 points at launch. Today it trails by 10.2.

Which AI Lab Releases New Models the Fastest?

Anthropic, by a wide margin. Its median gap between premium releases fell from 133 days in 2023 to 22.5 days in 2026, across ten releases in its Opus, Sonnet and Fable lines. Every other leading lab sits between six weeks and five months.

Lab20232024202520262026 releases
Anthropic1331087522.510
Alibaba (Qwen)-758345.56
MiniMax--142.5513
OpenAI118.56956.552.58
Zhipu AI (Z.ai)-2268356.54
xAI-121131595
Meta-136120604
Google DeepMind2105843.579.54
Moonshot AI--117833
DeepSeek-55.563127.52
Mistral AI751171121471

Median days between a lab's premium releases, counted in release days, so same-day siblings like GPT-5.4 and GPT-5.4 Pro do not create zero-day gaps.

OpenAI is the steady one, between 52 and 57 days every year since 2025. Google and DeepSeek slowed down in 2026. Google's Pro line has sat at Gemini 3.1 Pro since February, a teased Gemini 3.5 Pro never reached the API changelog, and Gemini 4 Argon is still limited to trusted testers. DeepSeek shipped twice in 2026, 111 days apart.

AI Model Release Timeline: Every Premium Release Since 2023

Every premium release from 15 labs, placed on its first day of public use in an app or API. The color on each model is the verdict against the model it replaced, and the number is its LMArena score change. Hover or tap a model for the predecessor, the win rate and the source. The impact chart under it puts every step since 2023 on one axis.

Gray models were never rated against their predecessor on LMArena, mostly 2023 and 2024 releases and gated models. Next release windows are estimates: each lab's last premium release plus the median of its last five gaps. They are not announced dates.

Which Releases Actually Moved the Needle

The same releases on one time axis. Bigger dots won more blind votes against the model they replaced, and color shows the verdict. Click a verdict to hide it, or narrow the view to 2025 or 2026.

Hollow dots are releases where one of the two models never appeared on LMArena, mostly 2023 and 2024 models and gated releases. LMArena's public leaderboard data starts in August 2024.

The full ledger runs from the first premium releases of 2023 to the latest AI models in this edition. Every chip carries the lab, release date, access route, predecessor and source. Hover or tap one to see them.

Should You Upgrade? Compare Any Two AI Models

Pick the model you use now and the one you are considering. The comparison uses blind head-to-head votes from LMArena, not launch claims. For direct successors you also get verified notes on what changed in daily use - speed, price, token use, hallucination and behavior - labelled by evidence type.

Read the verdict for what it is. LMArena measures which answer people prefer in open-ended chat. A model can lose there and still be the better pick for agentic coding or long tool chains, which is why the Intelligence Index change sits next to it wherever a like-for-like comparison exists.

Application

Stop re-testing every release yourself.

A new premium model lands every few days, and the newest is not always the right one for your task. Suprmind puts five frontier AIs in one conversation, and new models reach its AI Teams within days of release. When one model regresses on your question, the others read its answer and say so before you act on it.

See how the multi-AI workflow works

See the Multi-AI Platform in Action

Disagreement is the feature.

GPT-5.2 vs GPT-5.1: Why Fast Follow-Ups Disappoint

Point releases and snapshots that ship within 45 days of the previous version gained a median of +0.1 points. 5 of the 11 lost, and every one of those losses is statistically real. GPT-5.2 arrived 29 days after GPT-5.1 and wins 47.3% of votes against it. Claude Opus 4.8, Grok 4.6, Grok 4.7 and Qwen3-Max show the same pattern.

Fast is not always bad. GPT-6.1 Sol shipped 7 days after GPT-6 Sol and gained 26 points, mostly by fixing what GPT-6 Sol lost against its own predecessor. MiniMax M2.7 gained 24.1 after 34 days.

The time since the last release explains more than the version label. Compare releases that arrived after the same wait, and the label stops mattering:

Gap since previous versionNew generationPoint releaseSnapshot
46 to 90 days+13.2 (7 pairs)+10.8 (19 pairs)+10.2 (4 pairs)

Median LMArena change against the predecessor. Generations look bigger overall mainly because they arrive after longer waits.

Launch Benchmarks vs Blind Votes

Launch posts lead with the benchmark that moved most. Across 19 version pairs measured both ways, the median of the biggest gain each lab chose to report was +16.5 points. The same pairs moved +3 on the Artificial Analysis Intelligence Index. The two numbers sit on different scales, so read this as which number the launch post picked, not as a ratio.

  • GPT-5.2: +35.3 points on ARC-AGI-2 in the launch post. 18.5 LMArena points below GPT-5.1. [3]
  • Claude Opus 4.8: +27.4 points on USAMO at launch. 19.8 LMArena points below Claude Opus 4.7. [4]
  • DeepSeek-V4 Pro: +49.9 points on DeepSWE at launch. +9.3 on LMArena. [5]

None of this makes the launch numbers false. A model can get far better at competition math or agentic coding and feel no better, or worse, in conversation. The practical point: test a new version on your own work before you switch.

The newest model is not always the better one.
Let five of them check each other.

Versions ship faster than anyone can re-test them. In Suprmind, five frontier AIs answer in the same thread and read each other's work, so when the newest model gets your question wrong, another one catches it before you act on it.

Test It on a Real Question

7 days free. No credit card. Grok, GPT, Claude and Gemini in the trial.

When Do AI Labs Release New Models?

Monday to Thursday, almost always. Of 113 premium releases since January 2025, Thursday leads with 28, Friday has 8 and weekends 5. The busiest week on record was August 31 to September 6, 2026: Claude Fable 5.1, Gemini 3.8 Flash, Muse Spark 1.3, GPT-6 Astra and a Qwen3.8-Max snapshot shipped within three days.

Premium releases by weekday, January 2025 to October 3, 2026.

The theory that labs time launches to counter each other does not hold up. 103 of the 113 premium releases since January 2025, or 91%, landed within seven days of another lab's release. Spread the same releases over random dates and the share is still 91%.

Announced is not shipped. Google has the longest waits between announcing a model and letting the public use it: 73 days for Gemini 2.5 Deep Think, 64 for Gemini 1.0 Ultra and 54 for Gemini 1.5 Pro. This index dates every model by first public use, which is why Gemini 4 Argon is not in the ledger yet.

Can AI Research AI? How Five Deep Research Tools Did

We gave the same research brief to five deep research AIs and scored every answer against primary sources. The results are the Confidence Trap in fresh form. Two platforms called a 2026 slowdown that never happened, because each missed 22 or more of the year's 53 releases. One produced LMArena figures that match no snapshot LMArena has ever published. Only 49 of 195 models appeared in all five ledgers.

The AIs also know their own makers unevenly. ChatGPT and Claude got every release date for OpenAI and Anthropic exactly right. Grok got xAI's dates exactly right half the time, against 92% for other labs. Gemini found 11 of Google's 21 premium releases.

Disclosure. Claude scored this, including Claude's own run. The rules were fixed before scoring, and the scoring is mechanical.

When one AI regresses, another catches it.

Five frontier models in one thread, each reading the answers before it. Test it on a question where being wrong costs you.

Try Multi-AI Validation

Methodology and Data

What counts

  • Release date is the first public use in an app (paid tiers count) or an API. Waitlists, trusted-tester programs and arena-only testing do not count.
  • Premium means the lab positioned the model as its most capable at release, or it is a new version of a flagship line such as Opus, Sonnet, Pro, Sol or Muse Spark. Flash, mini and coding-only models count only when the lab called them its best overall.
  • One model, one entry. The same weights under two names count once. A general-availability release after a public preview of the same version counts as a snapshot.
  • Scope. 15 labs that had a model in the LMArena text top 10 or the OpenRouter token-share top 10 at any point since 2023. Baidu also meets that bar, since ERNIE 5.0 reached #9 on LMArena in three snapshots in December 2025 and January 2026, but it is not in the ledger yet. Headline figures use the 15.

How improvement is measured

  • Improvement and regression rates: the share of measured releases in a year rated a big or real step, or a regression. A regression is any loss that clears the 95% confidence band. An improvement needs a gap that clears the band and at least 51% of blind votes: a real step wins 51 to 55%, a big step 55% or more. A statistically real gain under 51% counts as a coin flip, which applies to one 2026 release, GPT-5.5 against GPT-5.4. Releases where either model never appeared on LMArena are left out of both rates.
  • Preference change: LMArena text leaderboard with style control, snapshot of Oct 2, 2026, highest-rated variant of each model. Win rate is estimated from the rating gap. A change counts as real when it clears the 95% confidence band. [1]
  • Capability change: Artificial Analysis Intelligence Index, only where both models were tested under the same index version. [2]
  • Usage notes: labelled as measured, user reports or lab claim. User reports need three independent sources. Lab claims never decide a verdict.

What we don't claim

LMArena preference is not task performance, and a statistical tie at the top is a tie. Dates for Alibaba, ByteDance and Moonshot rest partly on secondary sources. Silent updates behind the same model alias are not counted. Gemini experimental checkpoints never released to the public are not in the ledger.

Press kit and citation

You can quote any number or chart on this page with a link to the AI Models Index. Capability figures credit Artificial Analysis. Charts are free to use under CC BY 4.0. Every chart has Download, Embed and Cite buttons, and embedded charts update themselves with each edition.

Cite it as: Suprmind Data Lab, AI Models Index, October 4, 2026 edition, with a link to this page.

Get the press kit

High-resolution charts in light and dark, a fact sheet with the key numbers and how each one is measured, the method on one page, logos and a founder quote. The link appears as soon as you submit, and a copy goes to your inbox.

    Questions or a custom cut of the data: [email protected]. We handle your details under our privacy policy.

    Frequently Asked Questions About New AI Model Releases

    How often are new AI models released?

    In 2026 a premium AI model ships somewhere every 4.5 days, counting 15 leading labs. That is up from every 7.5 days in 2023. Individual labs are slower: a typical leading lab ships a premium model every six to twelve weeks, and Anthropic is the fastest at a median of 22.5 days in 2026.

    Which AI lab releases new models the fastest?

    Anthropic. Its median gap between premium releases was 22.5 days in 2026, across ten releases in the Opus, Sonnet and Fable lines. Alibaba's Qwen team was next at 45.5 days, then MiniMax at 51 and OpenAI at 52.5. Google and DeepSeek slowed down in 2026.

    Are new AI models actually better than the previous version?

    Usually, but by less each year. The median new release beat its predecessor by 21.5 LMArena points in 2025 and by 9.3 in 2026. Of 43 measured releases in 2026, 10 scored below the version they replaced, 7 of them by a statistically real margin. Capability benchmarks still rise: none of 39 version pairs went backwards on the Artificial Analysis Intelligence Index.

    Is GPT-5.2 better than GPT-5.1?

    Not for open-ended chat. In LMArena's blind head-to-head votes, GPT-5.2 wins 47.3% against GPT-5.1, an 18.5-point gap that is statistically real. OpenAI's launch post reported large gains on reasoning benchmarks such as ARC-AGI-2, so GPT-5.2 can still be the better pick for some tasks. Test it on your own work before you switch.

    What is the best AI model right now?

    There is no single answer. On LMArena's text leaderboard of October 2, 2026, Gemini 4 Argon holds #1 but is limited to trusted testers. The top three public models, Claude Opus 4.6, Claude Fable 5 and Claude Opus 5.5, sit within one rating point of each other, a statistical tie, and no model leads every task.

    How do you measure whether a new AI model is better?

    Each release is compared with the model it replaced: usually the previous version in the same model line, otherwise the lab's previous flagship. The comparison uses LMArena's text leaderboard with style control and the highest-rated variant of each model. The rating gap converts into the share of blind votes the newer model would win. A loss counts when it clears the 95% confidence band, and a gain when it also wins at least 51% of votes. Where a like-for-like comparison exists, the Artificial Analysis Intelligence Index adds a capability view.

    What counts as a premium AI model release?

    A model the lab positioned as its most capable at release, or a new version of a flagship line such as Claude Opus, Gemini Pro or GPT. The release date is the first day the public could use it in an app or API. Waitlists, trusted-tester programs and arena-only testing do not count.

    How often is the AI Models Index updated?

    A new edition ships every two weeks. Each edition adds new releases, refreshes the LMArena data, measures new version pairs and re-checks dates against lab announcements and API changelogs.

    Can I use this data?

    Yes, with attribution. Quote any number or chart with a link to the AI Models Index, and credit Artificial Analysis for Intelligence Index figures. Charts are free to use under CC BY 4.0. Every chart has Download, Embed and Cite buttons, and embedded charts update themselves with each edition. For the press kit with high-resolution charts and a fact sheet, use the press kit form or email [email protected].

    References

    1. LMArena. Leaderboard dataset, text leaderboard with style control, 187 snapshots from August 28, 2024 to October 2, 2026. huggingface.co/datasets/lmarena-ai/leaderboard-dataset
    2. Artificial Analysis. Intelligence Index and model comparisons. Figures used only where both models were tested under the same index version. artificialanalysis.ai
    3. OpenAI. "Introducing GPT-5.2." December 11, 2025. ARC-AGI-2 (Verified) 17.6% to 52.9%. openai.com/index/introducing-gpt-5-2
    4. Anthropic. "Claude Opus 4.8." May 28, 2026. www.anthropic.com/news/claude-opus-4-8
    5. DeepSeek. DeepSeek-V4 Pro release notes. August 13, 2026. api-docs.deepseek.com/news/news260813
    6. Suprmind Data Lab. "Multi-Model AI Divergence Index - The AI Confidence Trap." April 2026. suprmind.ai/hub/multi-model-ai-divergence-index
    7. OpenAI. API changelog (GPT-6 Sol, GPT-6.1 Sol release dates). developers.openai.com/api/docs/changelog
    8. Anthropic. "Claude Fable 5 and Claude Mythos 5." June 9, 2026. www.anthropic.com/news/claude-fable-5-mythos-5
    9. xAI. Developer release notes (Grok 4.5 general availability, July 8, 2026). docs.x.ai/developers/release-notes
    10. xAI. Grok release notes, April 17, 2026 (Grok 4.3). grok.com/release-notes/apr-17-2026
    11. Meta. "Introducing Muse Spark." April 8, 2026. about.fb.com/news/2026/04/introducing-muse-spark
    12. Moonshot AI. "Kimi K2.6." April 20, 2026. www.kimi.com/blog/kimi-k2-6
    13. Z.ai. "GLM-5." February 12, 2026. z.ai/blog/glm-5
    14. Mistral AI. "Mistral Medium 3." May 7, 2025. mistral.ai/news/mistral-medium-3
    15. Blog du Modérateur. LMArena top 10 language models, April 2024 (Command R+ and Mistral Large ranks, used for lab grouping). www.blogdumoderateur.com/ia-10-modeles-langage-performants-avril-2024

    Stop betting on one model's newest version.

    Five frontier models. One conversation. Every answer read and challenged by the next model in the thread. When a release disappoints, you find out on your question, not in next month's leaderboard.

    Compare Plans