GPT-5.6 Terra
GPT-5.6 Terra is an AI model released by OpenAI on Friday, Jun 26 2026, the same day as GPT-5.6 Sol. Benchmark results (shown below) cover BullshitBench v2, Gray Swan IPI, CursorBench v3.2, Frontier-Bench v0.1, Terminal-Bench 2.1, Arena Elo (Text), and 1 more.
Benchmarks
Nonsense detection
BullshitBench v2Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better.
53%
#18 of 67Best: Claude Opus 4.8 · 95%
Prompt injection robustness
Gray Swan IPIAttackers hide malicious instructions inside content the AI reads — a web page, an email, a document — and try to hijack what it does. Gray Swan's indirect prompt injection benchmark measures how often such an attack succeeds when the attacker gets a single try. Lower is better.
5.4%
k = 1
26%
k = 10
30.4%
k = 15
k = 1
#8 of 13Best: Claude Opus 5 · 0.2%
k = 10
#8 of 13Best: Claude Opus 5 · 1.6%
k = 15
#8 of 13Best: Claude Opus 5 · 2%
Agentic coding
CursorBench v3.2Cursor's own test of harder, real-world coding tasks inside a code editor, on the refreshed v3.2 task set. Scores aren't comparable with v3.1. Higher is better.
64.9%
#5 of 14Best: Claude Fable 5 · 70.5%
Agentic computer work
Frontier-Bench v0.1A hard, ever-evolving set of real computer tasks — coding, system administration, data work, and more — that an AI agent has to complete on its own. Run by the Harbor / Laude Institute team as the successor to Terminal-Bench (v0.1 is the first release of the task set). The score is the share of tasks solved. Higher is better.
20.8%
#5 of 9Best: Claude Opus 5 · 43.3%
Agentic terminal coding
Terminal-Bench 2.1Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better.
84.3%
#5 of 22Best: GPT-5.6 Sol · 88.8%
Community preference
Arena Elo (Text)Real people chat with two anonymous AIs side by side and vote for the answer they prefer. Votes become a chess-style Elo rating on arena.ai — it measures which AI people actually like, not test scores. Higher is better.
1467
#25 of 31Best: Claude Fable 5 · 1509
Community preference (code)
Arena Elo (Code)Like the text arena, but people vote on which AI writes better code. The votes become a chess-style Elo rating on arena.ai. Higher is better.
1526
#16 of 47Best: Kimi K3 · 1679
GPT-5.6 Terra — frequently asked questions
- When was GPT-5.6 Terra released?
- GPT-5.6 Terra was released by OpenAI on Friday, Jun 26 2026.
- Who made GPT-5.6 Terra?
- GPT-5.6 Terra was built by OpenAI. Creators of ChatGPT and the GPT series of models. Pioneered large-scale language model research.
- What benchmark scores did GPT-5.6 Terra get?
- GPT-5.6 Terra reports 9 tracked benchmark scores — BullshitBench v2: 53%; Gray Swan IPI (k = 1): 5.4%; Gray Swan IPI (k = 10): 26%; Gray Swan IPI (k = 15): 30.4%; CursorBench v3.2: 64.9%; Frontier-Bench v0.1: 20.8%; Terminal-Bench 2.1: 84.3%; Arena Elo (Text): 1467; Arena Elo (Code): 1526. Scores are the figures published at release by OpenAI.
- Is GPT-5.6 Terra open source?
- No. GPT-5.6 Terra is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.
- What came before and after GPT-5.6 Terra?
- OpenAI's previous tracked release was GPT-5.6 Sol on Jun 26 2026. It was followed by GPT-5.6 Luna on Jun 26 2026.
Compare GPT-5.6 Terra with
GPT-5.6 TerravsGPT-5.6-CyberGPT-5.6 TerravsClaude Opus 5GPT-5.6 TerravsGemini 3.6 FlashGPT-5.6 TerravsMuse GlimmerGPT-5.6 TerravsGrok 4.5GPT-5.6 TerravsDeepSeek-V4-Flash-0731GPT-5.6 TerravsMistral Medium 3.5GPT-5.6 TerravsKimi K3GPT-5.6 TerravsComposer 2.5GPT-5.6 TerravsGLM-5.2GPT-5.6 TerravsQwen3.8-Max