The eval you actually need syllabus
The average task is not your task
A benchmark is a sample of somebody else's distribution. Your product sees a narrower, messier one: repeated users, missing context, awkward images, and requests whose correct answer is refusal. The published score is useful as a prior, never as your product decision.
Ask what the score sampled
Before acting on a leaderboard, write down two things: what inputs were sampled and what counted as correct. If neither matches your product, the score is context rather than evidence. A three-point lead can disappear—or reverse—on the slice you actually serve.
Collect ten awkward inputs
Take ten real or realistically reconstructed inputs from your task. Prefer edge cases, ambiguous requests, missing information, and the examples people quietly remove from demos. Remove personal data before they become fixtures.
What does a leaderboard lead prove?
Model A leads a general benchmark by three points. What can you conclude about your production task?
Record the benchmark-to-task gap
Run the same candidate on your first probe set. Record the absolute point difference between its relevant published score and your task score. Keep the sign in your notes; this field captures the size of the mismatch.
Explain the inference limit
Explain to a teammate why a leaderboard is still useful even though it cannot make the product decision.