ML engineering
MLE-Bench
Can the AI do the work of a machine-learning engineer? It competes in real Kaggle competitions — building, training, and tuning models end to end — and the score reflects how well it places. Higher is better.
Rankings
Higher is betterMLE-Bench — frequently asked questions
- What is MLE-Bench?
- Can the AI do the work of a machine-learning engineer? It competes in real Kaggle competitions — building, training, and tuning models end to end — and the score reflects how well it places. Higher is better.
- Which AI model scores highest on MLE-Bench?
- Gemini 3.6 Flash by Google holds the best MLE-Bench result among tracked models, at 63.9% (released Jul 21 2026). Higher scores are better on this benchmark.
- What are the top 3 models on MLE-Bench?
- 1. Gemini 3.6 Flash (Google) — 63.9%; 2. Gemini 3.5 Flash (Google) — 49.7%; 3. Gemini 3.1 Pro (Google) — 42.6%.
- How many models have a published MLE-Bench score?
- 3 tracked models have a published MLE-Bench score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.