About
AI Release Tracker is one timeline of every notable model release from all the major AI labs. It records who shipped what, when, and how it scored at the time.
Why it started
The site went up on November 30, 2025: three years to the day after ChatGPT was released and set this pace. By that anniversary the labs were shipping faster than anyone could follow, and Peter Assentorp had a bad case of release FOMO. A flagship would land on a Tuesday, the discourse would move on by Thursday, and you would hear about it a week later from a screenshot.
The fix was a single page that answers the question directly: what came out, from whom, and how does it compare to the thing it replaced, with no feed to scroll through. That page became this site.
Where the numbers come from
Most scores are the ones the labs published themselves, at launch, and are recorded here as of that date rather than restated as current standings. Where a score comes from a leaderboard instead, the leaderboard is named — not only here, but beside that score on the model page, in the ranking on its benchmark page, in the Markdown twin of both, and in the JSON anyone can download. A gathered figure should never be able to travel on as a first-party one:
- Program reconstruction results from ProgramBench, backed by its MIT-licensed submissions registry. We report the percentage of fully solved programs using each model’s best published mini-SWE-agent result. Copyright Meta Platforms, Inc. and affiliates; MIT licence notice. Research by John Yang and colleagues (2026).
- Aggregated benchmark results from BenchLM, read from the dataset they publish themselves and used under the MIT licence they publish it with. They are where several of the coding and agentic figures are filled in when a lab published none.
- Terminal-Bench results from the benchmark’s own leaderboard, read from the submission files published in its Apache-2.0 repository. Each row is one agent, harness and reasoning effort, and the credit beside the score says which.
- Nonsense detection from BullshitBench, agent results from Supabase Evals and Next.js Evals.
- Community preference on Three.js scenes from threejseval, read from the JSON its ranking page is drawn from and used under the CC BY 4.0 it states for its vote and ranking data. The board ranks reasoning-effort variants separately; each release here carries the Elo of its best listed one.
- Who serves each model, and at what price, from OpenRouter.
Every one of them is worth reading at the source, where the methodology and the full field live. If a number here disagrees with theirs, theirs is right and we want to hear about it.
What is not here matters too. Everything above comes from a feed, a file or a page its publisher offers for the purpose. Where a source’s own terms say no, we do not take the data and we do not work around the answer — which is why the Arena Elo boards and Cursor’s CursorBench came off the site rather than staying on it. If you publish something we track and would rather we did not, tell us and it comes down.
Get in touch
The contact form is the fastest way to reach us, whatever the reason:
- Report a problem: a missing release, a benchmark score that looks wrong, a page that misbehaves
- Feature idea: something you want the site to do
- Sponsorship: placements and partnerships
- Something else: questions, press, or a note about the data