Evaluate before you ship
- prompting
- verification
- evaluation
Fluency is not quality. The better an output reads, the less you notice what is missing — a hallucination does not announce itself. A rubric is simply the questions you ask every time, in the same order, before anything leaves your hands. Five questions do most of the work.
The five questions
- Does it answer the actual ask? The deliverable you specified — format, length, audience. Not a similar topic: the ask.
- Which claims can you verify — and which can’t you? Numbers, names, quotes, dates, policies: trace them or mark them. Grounded, quote-first answers make this a ten-second job.
- What is missing that the reader would need? The gap between “covers the ask” and “usable” hides here.
- Would you put your name on it? Advice it cannot give, promises you cannot keep, tone that does not fit — this question catches them.
- Do the details fit the destination? If it will be pasted into a sheet or a tool, does it actually paste cleanly?
Comparing two drafts
When you have two versions, run the same rubric on both — question by question, not vibe against vibe. The more fluent one often wins the glance and loses the check; polish is not on the list. Score each question, then decide. This is the mirror image of why AI makes things up: plausible is generated, true is not, and the rubric is where the difference gets caught.
Scale the check to the stakes
For a throwaway draft, a skim is fine. For anything that goes to someone else — especially about people, money, policies, law or health — every claim gets traced, and unverifiable material gets cut or marked. You cannot delegate the last check: the model has no idea which details matter to your reader. That judgment is the job you are keeping.
A habit, not a ceremony
The five questions take about thirty seconds on a short text. Run them out loud once and they become automatic — the moment you catch yourself thinking “that’s a nice number, where did it come from?”, the rubric is doing its work.
The checklist survives its own use, too: when a rubric question never catches anything — or keeps catching everything — the question is wrong, so rewrite it. A rubric is a working tool, not a ceremony; four useful questions beat ten decorative ones.
Practice
Compare two replies and pick the one to ship
A colleague asked for a plain-language read of one clause, in three bullets, for a meeting in ten minutes. The excerpt says: 'Either party may terminate this Agreement for convenience upon written notice. Termination is effective upon receipt of the notice. Notice must be delivered to the other party's registered address.' You prompted an assistant and have two candidate replies — there is time to check one before the meeting.
Work through the rubric for both outputs, then pick the reply you would send.
Output A
Summary of the termination clause: - Either party may terminate for convenience with 15 days' written notice sent by email. - The contract renews automatically for five years unless terminated. This strongly favours the supplier. You should push to shorten the renewal term and secure a shorter notice period before signing.
Output B
Here is what the excerpt says about ending the contract early: - Either party can end the agreement for convenience — no reason needed. - Notice must be in writing; the excerpt does not state how many days. - It does not say whether notice by email is accepted — check the full contract. The excerpt does not mention automatic renewal, so I cannot confirm anything about it from this text.
Rubric — answer each question for both outputs before deciding
- Does it deliver the requested format — three bullets, plain language, no legalese?
- Can every specific claim (numbers, channels, lengths) be traced to the excerpt?
- Does it say what the excerpt does not cover instead of filling the gap?
- Does it stay a summary, rather than giving advice the writer cannot give?
Hint
Check every number in both replies against the excerpt — which reply contains numbers the excerpt never gives?
Which reply tells you what the source does not cover?
Why this is the answer
The rubric is the whole exercise. The more fluent reply fails four questions in a row — invented notice period, invented channel, invented renewal term, plus advice — while the plainer reply survives all four by staying inside the excerpt and naming its gaps. Polish is not a criterion; traceability is.
Transfer
- Before forwarding: questions 1 and 2 at minimum.
- Before publishing: all five, quote-first on every specific claim.
- On your own writing: the same five questions work — that is not an accident.
Next
That closes the prompting track — all eleven lessons. The research track is up next, starting with Search vs generation — when a question needs live sources before you trust the answer. Keep the muscles alive on the practice floor meanwhile.