promptengineering.ai

Model data

Which model, for which task — tested, not opined.

A fixed battery of everyday tasks, re-run on every notable model release. Each verdict is dated, shows its raw outputs, and links its scoring rubric.

  • Marketing copyClaude takes the lead — cleaner hooks, less boilerplate on 4 of 5 taskschanged
  • Long-doc summaryno change — GPT holds for the third release runningholds
  • Data extractionGemini, notably improved — if you gave up on it in spring, retry+18%

¹The full twenty-task battery, per-task pages, and raw outputs publish with the first public battery run. What you see here is the format, with sample data.