Knowledge work

AA-Briefcase

Artificial Analysis agentic office-work eval (Elo)

Rankings

Higher is better

AA-Briefcase — frequently asked questions

What is AA-Briefcase?
Artificial Analysis agentic office-work eval (Elo)
Which AI model scores highest on AA-Briefcase?
Grok 4.6 by SpaceXAI holds the best AA-Briefcase result among tracked models, at 1577 (released Aug 12 2026). Higher scores are better on this benchmark.
What are the top 2 models on AA-Briefcase?
1. Grok 4.6 (SpaceXAI) — 1577; 2. Grok 4.5 (SpaceXAI) — 1313.
How many models have a published AA-Briefcase score?
2 tracked models have a published AA-Briefcase score. Scores are the figures reported by each lab at that model's release, so this page is a record of results over time rather than a re-run leaderboard.
← All benchmarks