# Claude 3.7 Sonnet

Claude 3.7 Sonnet is an AI model released by Anthropic on Feb 24 2025. At release it scored 49% on BullshitBench v2, 68% on GPQA Diamond and 62.3% on SWE-Bench Verified.

## Facts

| Field | Value |
| --- | --- |
| Model | Claude 3.7 Sonnet |
| Developer | Anthropic |
| Release date | Monday, Feb 24 2025 |
| Licensing | Proprietary |

## Benchmark scores published at release

| Benchmark | Score | Source | What it measures |
| --- | --- | --- | --- |
| BullshitBench v2 | 49% | Lab | Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better. |
| SWE-Bench Verified | 62.3% | Lab | Real coding tasks pulled from open-source projects — the AI has to find and fix actual bugs. A human-checked version of the original SWE-Bench. Higher is better. |
| GPQA Diamond | 68% | Lab | Graduate-level science questions in biology, physics, and chemistry — hard enough that subject-matter PhDs score around 65%. Higher is better. |

## About Claude 3.7 Sonnet

Claude 3.7 Sonnet, released February 24, 2025, was Anthropic's first hybrid reasoning model: a single model that could answer instantly or engage an extended-thinking mode where it reasons step by step before responding, with the thinking budget under API control. It launched alongside the first research preview of Claude Code, Anthropic's agentic command-line coding tool.

The reasoning upgrade showed up most clearly in software engineering, where Claude 3.7 Sonnet scored 62.3% on SWE-Bench Verified — up from 49.0% for its predecessor and the best published score of any model at the time. GPQA Diamond rose to 68.0%. It held the top of the Claude lineup for three months until Claude Sonnet 4 and Opus 4 arrived in May 2025.

## Questions and answers

### When was Claude 3.7 Sonnet released?

Claude 3.7 Sonnet was released by Anthropic on Monday, Feb 24 2025.

### Who made Claude 3.7 Sonnet?

Claude 3.7 Sonnet was built by Anthropic. AI safety company building the Claude family of models. Founded in 2021 by former OpenAI researchers.

### What benchmark scores did Claude 3.7 Sonnet get?

Claude 3.7 Sonnet reports 3 tracked benchmark scores — BullshitBench v2: 49%; SWE-Bench Verified: 62.3%; GPQA Diamond: 68%. Scores are the figures published at release by Anthropic.

### Is Claude 3.7 Sonnet open source?

No. Claude 3.7 Sonnet is a proprietary model. The weights are not published — it is available only through the provider's own API, apps, or partner platforms.

### What came before and after Claude 3.7 Sonnet?

Anthropic's previous tracked release was Claude 3.5 Sonnet (upgraded) on Oct 22 2024, 125 days earlier. It was followed by Claude Sonnet 4 on May 22 2025.


---

Canonical page: https://aireleasetracker.com/model/anthropic/claude-3.7-sonnet
Full dataset: https://aireleasetracker.com/llms-full.txt · JSON: https://aireleasetracker.com/models.json
Source: AI Release Tracker (https://aireleasetracker.com). Benchmark scores are the figures published by the releasing lab at launch.
