Long-horizon office work

CoWorkBench

Tests long-running office tasks across fields including computer science, finance, law, medicine, and other productivity work. Higher is better.

Rankings

Higher is better

CoWorkBench — frequently asked questions

What is CoWorkBench?
Tests long-running office tasks across fields including computer science, finance, law, medicine, and other productivity work. Higher is better.
Which AI model scores highest on CoWorkBench?
Qwen3.8-27B by Qwen holds the best CoWorkBench result among tracked models, at 70.7% (released Aug 14 2026). Higher scores are better on this benchmark.
What are the top 1 models on CoWorkBench?
1. Qwen3.8-27B (Qwen) — 70.7%.
How many models have a published CoWorkBench score?
1 tracked model has a published CoWorkBench score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
What is the best open model on CoWorkBench?
Qwen3.8-27B by Qwen is the highest-ranked model with downloadable weights on CoWorkBench, scoring 70.7% at rank 1 overall.
← All benchmarks