Supabase coding

Supabase Evalsno skills

The same Supabase scenarios, but with none of Supabase's skills loaded — so it measures what the model already knows about building on Supabase, rather than how well it follows Supabase's supplied instructions. Higher is better.

Rankings

Higher is better

Supabase Evals — frequently asked questions

What is Supabase Evals?
The same Supabase scenarios, but with none of Supabase's skills loaded — so it measures what the model already knows about building on Supabase, rather than how well it follows Supabase's supplied instructions. Higher is better.
Which AI model scores highest on Supabase Evals?
Kimi K3 by Moonshot AI holds the best Supabase Evals (no skills) result among tracked models, at 100% (released Jul 16 2026). Higher scores are better on this benchmark.
What are the top 5 models on Supabase Evals?
1. Kimi K3 (Moonshot AI) — 100%; 2. GPT-5.6 Sol (OpenAI) — 94.7%; 2. Claude Opus 5 (Anthropic) — 94.7%; 4. Claude Sonnet 5 (Anthropic) — 78.9%; 5. GPT-5.4 mini (OpenAI) — 73.7%.
How many models have a published Supabase Evals score?
5 tracked models have a published Supabase Evals (no skills) score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks