When you need to judge whether a Claude response is good, the fastest option is often to just ask Claude: paste the response back in and ask "is this accurate and complete?" It's tempting because it's immediate and requires no separate setup. It's also one of the least reliable ways to evaluate output, and the exam tests this directly because the failure mode is subtle — the grading response often sounds just as confident and well-reasoned as everything else Claude writes.
Why the Same Reasoning Grades Itself as Sound
The core problem isn't laziness or a lack of effort in the grading step. It's structural: the same reasoning patterns, knowledge gaps, and assumptions that produced the original output are what get applied when that output is evaluated. If a claim seemed well-supported enough to state confidently the first time, the same underlying reasoning will tend to find it well-supported the second time too, because nothing about the evaluation pass changes what the model actually knows or how it weighs evidence. A wrong assumption made while generating an answer doesn't become visible just because you asked the same instance to look again — it's not a fresh perspective, it's the same perspective checking its own work with the same blind spots intact.
This is easy to underestimate because a self-grading response can still look rigorous: it might list several criteria, note a minor caveat, and land on a confident score. The rigor of the grading format doesn't fix the underlying issue — the same source of error that produced a flawed claim is also the source doing the checking.
What a Better Evaluation Pattern Looks Like
A more reliable pattern has three ingredients, useful in combination or alone depending on the stakes:
- Independent, concrete criteria. A checklist of specific pass/fail items ("does the answer cite a specific source for each statistic?", "does the recommendation acknowledge at least one tradeoff?") gives an evaluator something objective to check, rather than an open-ended "is this good?" judgment that invites the same reasoning that produced the answer.
- A genuinely separate evaluation pass. A fresh conversation with an explicit critic framing ("you are reviewing this for factual errors and unsupported claims — be skeptical") is stronger than asking within the same thread, though it's worth being honest that it's still the same underlying model and won't catch everything a truly independent reviewer would.
- A human reviewer for anything consequential. For output that will be published, acted on, or affects a real decision, a human check is the most reliable independent signal available — nothing about a model, fresh conversation or not, replaces that for high-stakes cases (see lesson 2.5).
Key Concept
Self-grading is unreliable because the same reasoning that generated an error is the reasoning applied to check for it — a confident-sounding grade doesn't mean an independent check happened. Prefer explicit, concrete criteria; a genuinely separate evaluation pass; and human review for anything consequential.
Common Exam Distractor
Watch for an answer that "fixes" self-grading by making the grading instructions more detailed or asking the same conversation to role-play as a harsh critic. A stricter rubric or a critic persona applied by the same model, in the same context, still carries forward the reasoning and assumptions that produced the original output — it narrows the problem but doesn't remove its structural cause.