# Supabase Evals (with skills) — AI model rankings

Supabase's own open benchmark: a coding agent is dropped into a real Supabase project and asked to do real work — set up a schema, fix a broken security policy, debug an Edge Function — and every run is checked against a live Supabase stack. This is the headline number, where the agent has Supabase's own skills loaded, as most people building on Supabase would. The score is the share of scenarios it got right. Higher is better.

5 tracked models have a published Supabase Evals (with skills) score. Higher is better. Scores are as published at each model's release.

## Ranking

| Rank | Model | Developer | Score | Released |
| --- | --- | --- | --- | --- |
| 1 | GPT-5.6 Sol | OpenAI | 100% | Jun 26 2026 |
| 2 | Claude Sonnet 5 | Anthropic | 94.7% | Jun 30 2026 |
| 2 | Kimi K3 | Moonshot AI | 94.7% | Jul 16 2026 |
| 2 | Claude Opus 5 | Anthropic | 94.7% | Jul 24 2026 |
| 5 | GPT-5.4 mini | OpenAI | 78.9% | Mar 17 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/supabase-evals
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
