OpenAI
GPT-5.6 Sol
GPT-5.6 Sol is an AI model released by OpenAI on Friday, Jun 26 2026, 64 days after GPT-5.5-Pro. Benchmark results (shown below) cover BullshitBench v2, Next.js Evals, Terminal-Bench 2.1, GDPval-AA v2, Arena Elo (Text), and Arena Elo (Code).
Benchmarks
Nonsense detection
BullshitBench v2Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.
47%
Next.js coding
Next.js EvalsVercel's open eval of how well AI coding agents build and migrate real Next.js apps — measured as the share of tasks the agent completes successfully. Higher is better.
92%
Agentic terminal coding
Terminal-Bench 2.1Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better.
88.8%
Knowledge work
GDPval-AA v2economically valuable knowledge work (v2, re-based Elo)
1748
Community preference
Arena Elo (Text)Real people chat with two anonymous AIs side by side and vote for the answer they prefer. Votes become a chess-style Elo rating on arena.ai — it measures which AI people actually like, not test scores. Higher is better.
1486
Community preference (code)
Arena Elo (Code)Like the text arena, but people vote on which AI writes better code. The votes become a chess-style Elo rating on arena.ai. Higher is better.
1636