The eval you actually need syllabus
Three failures a benchmark will never show you
Aggregate accuracy compresses unlike failures into one number. A harmless formatting miss, a confident fabrication, and a dangerous refusal can all move the score by exactly one point while carrying completely different product costs.
Separate wrong from dangerous
Start with three buckets: near-miss (the intent is right but the form is wrong), confident nonsense (plausible and unsupported), and bad refusal behaviour (answers when it should abstain, or refuses when it should help). Add domain-specific buckets only when two failures demand different fixes.
Label a failure sample
Run at least twenty probes and assign every failure exactly one primary category. Keep the raw response beside the label so you can audit your own taxonomy later.
Why keep failure categories?
Two models both score 82%. Why might one still be the clear product choice?
Count the categories you actually need
Record how many non-empty primary failure categories remained after labelling. Fewer is not automatically better; the goal is a small set where every bucket leads to a distinct decision.
Triage a misleading aggregate
A new model raises aggregate accuracy but doubles unsafe answers on ambiguous inputs. The launch review is tomorrow. What do you do?