Supabase coding
Supabase Evalswith skills
Supabase's own open benchmark: a coding agent is dropped into a real Supabase project and asked to do real work — set up a schema, fix a broken security policy, debug an Edge Function — and every run is checked against a live Supabase stack. This is the headline number, where the agent has Supabase's own skills loaded, as most people building on Supabase would. The score is the share of scenarios it got right. Higher is better.
Rankings
Higher is betterSupabase Evals — frequently asked questions
- What is Supabase Evals?
- Supabase's own open benchmark: a coding agent is dropped into a real Supabase project and asked to do real work — set up a schema, fix a broken security policy, debug an Edge Function — and every run is checked against a live Supabase stack. This is the headline number, where the agent has Supabase's own skills loaded, as most people building on Supabase would. The score is the share of scenarios it got right. Higher is better.
- Which AI model scores highest on Supabase Evals?
- GPT-5.6 Sol by OpenAI holds the best Supabase Evals (with skills) result among tracked models, at 100% (released Jun 26 2026). Higher scores are better on this benchmark.
- What are the top 5 models on Supabase Evals?
- 1. GPT-5.6 Sol (OpenAI) — 100%; 2. Claude Sonnet 5 (Anthropic) — 94.7%; 2. Kimi K3 (Moonshot AI) — 94.7%; 2. Claude Opus 5 (Anthropic) — 94.7%; 5. GPT-5.4 mini (OpenAI) — 78.9%.
- How many models have a published Supabase Evals score?
- 5 tracked models have a published Supabase Evals (with skills) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.