# CheatBench — AI model rankings

Measures how often AI agents try to cheat on difficult assignments, such as reading hidden answers, copying work or manipulating grading. The overall score gives equal weight to ten categories; the sycophancy category measures how far an agent shifts its beliefs toward a user's stated views. Scores describe each model in its tested agent setup, and task success is measured separately. Lower is better.

Scores from [CheatBench (Center for AI Safety)](https://cheatbench.ai/), which runs the benchmark and publishes the full field.

12 tracked models have a published CheatBench score. Lower is better. Scores are gathered from CheatBench (Center for AI Safety); recorded retrieval dates appear beside the scores.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | Claude Opus 5.5 | Anthropic | 11.2% | [CheatBench](https://cheatbench.ai/), retrieved 2026-09-25 | Sep 22 2026 |
| 2 | Muse Spark 1.3 | Meta | 39% | [CheatBench](https://cheatbench.ai/), retrieved 2026-09-25 | Sep 2 2026 |
| 3 | Claude Opus 5 | Anthropic | 45.3% | [CheatBench](https://cheatbench.ai/), retrieved 2026-09-25 | Jul 24 2026 |
| 4 | Claude Fable 5.1 | Anthropic | 45.8% | [CheatBench](https://cheatbench.ai/), retrieved 2026-09-25 | Sep 1 2026 |
| 5 | GPT-6 Astra | OpenAI | 47.4% | [CheatBench](https://cheatbench.ai/), retrieved 2026-09-25 | Sep 3 2026 |
| 6 | Kimi K3 | Moonshot AI | 70% | [CheatBench](https://cheatbench.ai/), retrieved 2026-09-25 | Jul 16 2026 |
| 7 | GPT-6 Sol | OpenAI | 71.9% | [CheatBench](https://cheatbench.ai/), retrieved 2026-09-25 | Sep 22 2026 |
| 8 | DeepSeek-V4-Pro | DeepSeek | 75.1% | [CheatBench](https://cheatbench.ai/), retrieved 2026-09-25 | Apr 24 2026 |
| 9 | Gemini 3.8 Flash | Google | 75.2% | [CheatBench](https://cheatbench.ai/), retrieved 2026-09-25 | Sep 2 2026 |
| 10 | GPT-5.6 Sol | OpenAI | 77% | [CheatBench](https://cheatbench.ai/), retrieved 2026-09-25 | Jun 26 2026 |
| 11 | Grok 4.7 | SpaceXAI | 78% | [CheatBench](https://cheatbench.ai/), retrieved 2026-09-25 | Sep 21 2026 |
| 12 | Grok 4.6 | SpaceXAI | 79.1% | [CheatBench](https://cheatbench.ai/), retrieved 2026-09-25 | Aug 12 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/cheatbench
Site index: https://aireleasetracker.com/llms.txt
Source: AI Release Tracker (https://aireleasetracker.com). Most benchmark scores come from lab launch material; gathered results identify the leaderboard that published them.
