Model Watch
When a model ships, we re-test everything.
Within 48 hours of a notable release, the full task battery re-runs and each report answers one question: what changed for the work you actually do?
Specimen · a watch emailformat preview
Subject: Gemini 3.1 — 3 of 20 practitioner
tasks change verdict
Marketing copy: Gemini overtakes GPT
(examples inside).
Long-doc summarization: unchanged, Claude holds.
Data extraction: improved 18% — if you gave
up on Gemini, retry.
→ Your watched prompt ("LinkedIn post"):
re-test on the new lineup in one click.Keep this working
Get the next watch report.
Release-triggered only. No weekly filler — if nothing ships, nothing arrives.