Model data
Which model, for which task — tested, not opined.
A fixed battery of everyday tasks, re-run on every notable model release. Each verdict is dated, shows its raw outputs, and links its scoring rubric.
battery № 47 · run 2026-08-25
- Marketing copyClaude takes the lead — cleaner hooks, less boilerplate on 4 of 5 taskschanged
- Long-doc summaryno change — GPT holds for the third release runningholds
- Data extractionGemini, notably improved — if you gave up on it in spring, retry+18%
¹The full twenty-task battery, per-task pages, and raw outputs publish with the first public battery run. What you see here is the format, with sample data.