Here's the prompt ๐ฑ It takes any workflow you already run and tells you where your judgment actually belongs, then labels which of those checks you can hand back to Claude later.
one useful run
Paste this into Claude ๐ฉโ๐ป
Use it in the chat or project that already contains the task, so it has your real workflow to work from. If the material isn't there yet, paste or attach it right after the prompt.
copy this promptTurn the workflow in this conversation into an AI approval loop. Place human judgment after any pass where truth, safety, taste, intent or an irreversible action can diverge. Use objective tests for facts, files, calculations and system state; reserve human review for consequential tradeoffs and taste.
Return the workflow as: human sets outcome and standard โ AI pass โ check or approval โ AI revision โ final verification. For every checkpoint, state what evidence is required and what happens on failure. Then propose how repeated corrections become rules, examples, Skills or evals, and which single checkpoint could be removed first only after the workflow repeatedly passes.
๐ฑ What a good result looks like. Your actual workflow rewritten as a loop, each checkpoint labelled objective or human, with the evidence it needs and what happens when it fails. If it hands you a blank framework to fill in, say "do it for the workflow above, using my actual steps" and run it again.
reading the output
Objective check, or human call? โ
This is the label that decides everything downstream, so it's worth being able to check its work:
โ
Objective check There's a correct answer and something other than you could confirm it. Does the link resolve, does the number match the source, does the code run, is the claim in the cited document. These are the ones that graduate into scripts, tests and evals.
๐ง Human call There's no correct answer, only a defensible one. What to lead with, is this tone right for this person, does this sound like me. These never graduate. Automating them is how work becomes average.
The objective ones are your automation backlog. The human ones are the job. If Claude labels a taste decision as objective, move it back.
automating them later
How a checkpoint graduates โฌ๏ธ
Every time you correct the same thing twice, that correction moves up a rung:
- 1๏ธโฃ A rule Write it into the prompt or your project instructions. "Never open with a question." Cheapest possible fix.
- 2๏ธโฃ An example When a rule keeps getting misread, show instead of tell. One good and one bad, labelled. Examples beat adjectives.
- 3๏ธโฃ A skill When the same multi-step correction recurs, save the whole procedure so it runs the same way without you restating it.
- 4๏ธโฃ An eval A repeatable test that proves the thing is right, so the check runs without you reading it. Objective checkpoints only.
- 5๏ธโฃ Remove the gate Only after the eval has passed repeatedly on real work, and one gate at a time so you can tell what broke.
โ ๏ธ Remove the gate that has never caught anything, not the one that annoys you most. The annoying one is usually annoying because it keeps finding things.
what it looks like
Four workflows, gated ๐ฝ
Notice how few checkpoints there are, and how specific each one is:
โ๏ธ A published post Set the angle โ AI drafts โ check intent and taste โ AI tightens โ verify every claim and link. One taste gate, one truth gate.
๐ง Outreach email Set the ask โ AI drafts โ check tone and the ask โ AI revises โ you send it yourself. That last one is why draft-only is a rule, not a preference.
๐ Research summary Set the question โ AI gathers โ check the sources are real and say what it claims โ AI synthesises โ check nothing was invented in the compression. Both objective, so both automatable later.
๐ป A code change Set the behaviour โ AI writes โ tests run (already automated) โ AI fixes โ you review the approach. Judgment goes on design, the part tests can't see.
๐จ Some gates never come out. Health, legal, financial, security, privacy and anything irreversible keep a human, however well the workflow has performed. No prompt, skill or eval makes an agent injection-proof, and a long clean run is not evidence that it is.
Sources: the AI sandwich framing ยท the 758-consultant study ยท turning checks into evals ยท NIST on human oversight