# BrowseComp — AI model rankings

Can the AI browse the web and track down hard-to-find answers? Higher is better.

28 tracked models have a published BrowseComp score. Higher is better. Scores come from published lab reports and benchmark sources; recorded sources appear beside the scores.

## Ranking

| Rank | Model | Developer | Score | Source | Released |
| --- | --- | --- | --- | --- | --- |
| 1 | GPT-5.6 Sol | OpenAI | 92.2% | [BenchLM](https://benchlm.ai), retrieved 2026-08-24 | Jun 26 2026 |
| 2 | GPT-6 Astra | OpenAI | 91.5% | Lab | Sep 3 2026 |
| 3 | Kimi K3 | Moonshot AI | 91.2% | Lab | Jul 16 2026 |
| 4 | Claude Opus 5 | Anthropic | 90.8% | Lab | Jul 24 2026 |
| 5 | GPT-5.5-Pro | OpenAI | 90.1% | Lab | Apr 23 2026 |
| 6 | GPT-5.4-Pro | OpenAI | 89.3% | Lab | Mar 5 2026 |
| 7 | Claude Mythos 5 | Anthropic | 88% | [BenchLM](https://benchlm.ai), retrieved 2026-08-24 | Jun 9 2026 |
| 8 | GPT-5.6 Terra | OpenAI | 87.5% | [BenchLM](https://benchlm.ai), retrieved 2026-08-24 | Jun 26 2026 |
| 9 | Claude Fable 5 | Anthropic | 86.9% | [BenchLM](https://benchlm.ai), retrieved 2026-07-11 | Jun 9 2026 |
| 10 | Gemini 3.1 Pro | Google | 85.9% | Lab | Feb 19 2026 |
| 11 | Claude Sonnet 5 | Anthropic | 84.7% | [BenchLM](https://benchlm.ai), retrieved 2026-07-11 | Jun 30 2026 |
| 12 | GPT-5.5 | OpenAI | 84.4% | Lab | Apr 23 2026 |
| 13 | Claude Opus 4.8 | Anthropic | 84.3% | [BenchLM](https://benchlm.ai), retrieved 2026-07-11 | May 28 2026 |
| 14 | Claude Opus 4.6 | Anthropic | 83.7% | [BenchLM](https://benchlm.ai), retrieved 2026-07-11 | Feb 5 2026 |
| 15 | DeepSeek-V4-Pro | DeepSeek | 83.4% | [BenchLM](https://benchlm.ai), retrieved 2026-07-11 | Apr 24 2026 |
| 15 | DeepSeek-V4-Pro-0813 | DeepSeek | 83.4% | [BenchLM](https://benchlm.ai), retrieved 2026-08-31 | Aug 13 2026 |
| 17 | GPT-5.6 Luna | OpenAI | 83.3% | [BenchLM](https://benchlm.ai), retrieved 2026-08-31 | Jun 26 2026 |
| 18 | Kimi K2.6 | Moonshot AI | 83.2% | [BenchLM](https://benchlm.ai), retrieved 2026-08-31 | Apr 21 2026 |
| 19 | GPT-5.4 | OpenAI | 82.7% | Lab | Mar 5 2026 |
| 20 | Claude Opus 4.7 | Anthropic | 79.3% | Lab | Apr 16 2026 |
| 21 | GLM-5 | Z.ai | 75.9% | Lab | Feb 12 2026 |
| 22 | DeepSeek-V4-Flash-0731 | DeepSeek | 73.2% | [BenchLM](https://benchlm.ai), retrieved 2026-08-31 | Jul 31 2026 |
| 23 | Qwen3.5 | Qwen | 69% | Lab | Feb 16 2026 |
| 24 | GLM-5.1 | Z.ai | 68% | Lab | Apr 7 2026 |
| 25 | GPT-5.2 | OpenAI | 65.8% | [BenchLM](https://benchlm.ai), retrieved 2026-08-31 | Dec 11 2025 |
| 26 | Kimi K2.5 | Moonshot AI | 60.6% | [BenchLM](https://benchlm.ai), retrieved 2026-08-31 | Jan 27 2026 |
| 27 | GLM-4.7 | Z.ai | 52% | Lab | Dec 22 2025 |
| 28 | Nemotron 3 Ultra | NVIDIA | 44.4% | [BenchLM](https://benchlm.ai), retrieved 2026-08-31 | Jun 4 2026 |


---

Canonical page: https://aireleasetracker.com/benchmark/browsecomp
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Most benchmark scores come from lab launch material; gathered results identify the leaderboard that published them.
