Spatial reasoning
Blueprint-Bench 2
Can the AI reason about space and layout — for example, understanding a floor plan or blueprint? Higher is better.
Rankings
Higher is betterBlueprint-Bench 2 — frequently asked questions
- What is Blueprint-Bench 2?
- Can the AI reason about space and layout — for example, understanding a floor plan or blueprint? Higher is better.
- Which AI model scores highest on Blueprint-Bench 2?
- GPT-5.5 by OpenAI holds the best Blueprint-Bench 2 result among tracked models, at 36.2% (released Apr 23 2026). Higher scores are better on this benchmark.
- What are the top 5 models on Blueprint-Bench 2?
- 1. GPT-5.5 (OpenAI) — 36.2%; 2. Gemini 3.5 Flash (Google) — 33.6%; 3. Gemini 3.1 Pro (Google) — 26.5%; 4. Claude Opus 4.7 (Anthropic) — 24.5%; 5. Claude Sonnet 4.6 (Anthropic) — 6.7%.
- How many models have a published Blueprint-Bench 2 score?
- 6 tracked models have a published Blueprint-Bench 2 score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.