# Supabase Evals (with skills) — AI model rankings

Supabase's own open benchmark: a coding agent is dropped into a real Supabase project and asked to do real work — set up a schema, fix a broken security policy, debug an Edge Function — and every run is checked against a live Supabase stack. This is the headline number, where the agent has Supabase's own skills loaded, as most people building on Supabase would. The score is the share of scenarios it got right. Higher is better.

Scores from [Supabase Evals](https://supabase.com/evals), which runs the benchmark and publishes the full field.

6 tracked models have a published Supabase Evals (with skills) score. Higher is better. Scores are gathered from Supabase Evals; recorded retrieval dates appear beside the scores.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | Claude Sonnet 5 | Anthropic | 91.3% | [Supabase Evals](https://supabase.com/evals) | Jun 30 2026 |
| 1 | Claude Opus 5 | Anthropic | 91.3% | [Supabase Evals](https://supabase.com/evals) | Jul 24 2026 |
| 3 | GPT-5.6 Sol | OpenAI | 89.9% | [Supabase Evals](https://supabase.com/evals) | Jun 26 2026 |
| 4 | Kimi K3 | Moonshot AI | 84.1% | [Supabase Evals](https://supabase.com/evals) | Jul 16 2026 |
| 5 | GPT-5.6 Luna | OpenAI | 68.1% | [Supabase Evals](https://supabase.com/evals) | Jun 26 2026 |
| 6 | GPT-5.4 mini | OpenAI | 65.2% | [Supabase Evals](https://supabase.com/evals) | Mar 17 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/supabase-evals
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Most benchmark scores come from lab launch material; gathered results identify the leaderboard that published them.
