Coffee break: the summariser you can prove is wrong
Twenty minutes to build it, twenty more to find out where it fails. The second half is the point.
Forty minutes, two halves
The first twenty build a meeting summariser. That part is genuinely easy now and you have probably seen it done a dozen times.
The second twenty are the ones nobody films: five probes that catch it being confidently wrong. By the end you will have a number — how many of your five it passes — and that number is worth more than the summariser.
What you need
- Any model you can call, hosted or local. It does not matter which; the method is the point.
- One real transcript. A meeting, a call, a voice memo of yourself talking for five minutes.
- No account, no hardware, no framework.
Where this leads
If the second half interests you more than the first, that is the signal to take the eval course next. This build is a twenty-minute version of what that course does properly.
The coffee break
Build it, then break it.
- Twenty minutes to a summariser — The easy half, done deliberately plainly so there is nothing to hide behind later. (20 min)
- Twenty minutes to break it — Five probes, one number, and the reason you will never trust a demo again. (20 min)
Related writing
- You don't need a benchmark, you need five probes — Public evals measure the average task. Yours has awkward inputs and one definition of wrong. Building the small, specific eval that catches your failures in an afternoon.