Claude Mythos 5.1vsQwen3.8-Max
Claude Mythos 5.1 | Qwen3.8-Max | |
|---|---|---|
| Specifications | ||
ParametersA rough measure of how big the model is. More parameters usually means more capable and more expensive to run, though it is a poor guide on its own — a smaller, newer model often beats a larger, older one. | — | 2.4T |
Context windowHow much text the model can hold in mind at once — your question, any documents you attach, the conversation so far, and its own reply. Go past it and the earliest part falls out of view. | 1M | 1M |
| API pricingUSD per 1M tokens · lower wins | ||
Cheapest inputLowest input rate across third-party providers, excluding the lab itself. The cheapest endpoint may run a quantised build or a shorter context — see "Available from" on the model page. | — | $2.00Alibaba |
Cheapest outputLowest output rate across third-party providers, excluding the lab itself. May come from a different provider than the cheapest input. | — | $6.00Alibaba |
| BenchmarksPublished by one model only | ||
BullshitBench v2Nonsense detection — Given a confidently-worded but nonsensical prompt, does the AI spot that it makes no sense and push back — instead of playing along and inventing an answer? The score is how often it clearly called out the nonsense. Higher is better. | — | 94% |
SWE-Bench ProAgentic coding — Can the AI fix real bugs in real software? It's handed actual problems from open-source projects and has to write code that genuinely solves them. Higher is better. | — | 67.7% |
PaperBenchResearch reproduction — reproducing the results of an ML research paper end to end | — | 93% |
Terminal-Bench 4.0Agentic terminal coding — Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Version 4.0 recalibrated how much time, CPU and memory each task gets, removed eight tasks and fixed nineteen, so fewer runs fail for reasons that have nothing to do with the model. Scores are not comparable with earlier versions. Higher is better. | 60.9% | — |
Terminal-Bench 2.1Agentic terminal coding — Can the AI work in a command-line terminal — running commands and finishing technical setup tasks the way a developer would? Higher is better. | — | 86.6% |
JobBenchProfessional tool use — Tests the AI on professional workplace tasks that require using real work tools — the kind of multi-step jobs an office worker handles. Higher is better. | — | 53.4% |
Humanity's Last Exam · with toolsMultidisciplinary reasoning — Humanity's Last Exam — extremely hard expert questions across many subjects. “With tools” means the AI is allowed to search the web or run code while answering. Higher is better. | — | 43.6% |
OSWorld-VerifiedAgentic computer use — Can the AI actually operate a computer — clicking, typing, and using real apps — to finish tasks on its own? Higher is better. | — | 86.1% |
CharXiv ReasoningChart reasoning — Can the AI read and reason about complex charts and figures, not just text? Higher is better. | — | 88.4% |
BabyVisionVisual reasoning — Tests core visual reasoning — seeing and understanding images the way even young children can, which AIs often find surprisingly hard. Higher is better. | — | 82% |
| Overview | ||
| Company | Anthropic | Qwen |
| Release date | Sep 1 2026 | Aug 3 2026 |
| Access | Proprietary | Proprietary |
Other comparisons
Claude Mythos 5.1vsGPT-5.6 SolQwen3.8-MaxvsGPT-5.6 SolClaude Mythos 5.1vsGemini 3.7 FlashQwen3.8-MaxvsGemini 3.7 FlashClaude Mythos 5.1vsMuse GlimmerQwen3.8-MaxvsMuse GlimmerClaude Mythos 5.1vsGrok 4.6Qwen3.8-MaxvsGrok 4.6Claude Mythos 5.1vsDeepSeek-V4-Pro-0813Qwen3.8-MaxvsDeepSeek-V4-Pro-0813Claude Mythos 5.1vsMistral Medium 3.5Qwen3.8-MaxvsMistral Medium 3.5
Frequently asked questions
Claude Mythos 5.1 and Qwen3.8-Max don't publish scores on any of the same benchmarks, so there's no direct head-to-head comparison. Qwen3.8-Max shipped 29 days before Claude Mythos 5.1, so benchmark comparisons should account for the intervening progress.
Context windows are 1M (Claude Mythos 5.1) vs 1M (Qwen3.8-Max).
Direct benchmark comparisons are unavailable — Claude Mythos 5.1 and Qwen3.8-Max don't publish scores on any of the same benchmarks.
Claude Mythos 5.1 was released by Anthropic on Sep 1 2026.
Qwen3.8-Max was released by Qwen on Aug 3 2026.
Claude Mythos 5.1 has a 1M context window; Qwen3.8-Max has 1M.