Some decisions carry consequences serious enough for a real person — their job, their health, their legal standing, their access to credit — that full autonomy is inappropriate no matter how capable the underlying model is. Recognizing which categories of decision call for this, and knowing concretely what "human review required" actually looks like in practice, rather than as a vague caveat, is a distinct skill the exam tests directly.
Categories That Almost Always Need a Human in the Loop
Four categories come up repeatedly, each for a related but slightly different reason:
- Hiring. A resume-screening or candidate-ranking process built on Claude can reflect biases present in the input data — patterns correlated with protected characteristics can influence an output even without anyone intending that. A human needs to review the actual final decision, not just spot-check the process once at setup.
- Medical. Claude can produce medically plausible-sounding text that's still wrong, and a wrong medical claim delivered confidently is exactly the hallucination risk from lesson 2.1 at its highest stakes. Output here should be treated as a draft or a decision-support input for a licensed professional, never as the diagnosis or treatment decision itself.
- Legal. Claude can fabricate a plausible-sounding case citation, statute, or precedent (this is a well-documented failure mode). Anything that will be filed, cited, or relied on in an actual legal matter needs a qualified reviewer to independently confirm every citation before it's used, not just a read-through for tone.
- Financial decisions affecting individuals. Credit, lending, and eligibility decisions carry both a fairness risk (biased outcomes from data patterns) and, in many jurisdictions, an actual accountability requirement — a human needs to be able to explain and stand behind an adverse decision, which an unreviewed automated output can't do on its own.
What 'Additional Verification' Actually Means
"Requires human review" is easy to state and easy to reduce to a meaningless rubber stamp. A verification step that actually does something has three properties: it's performed by someone qualified to judge the specific content (a licensed clinician for medical content, an attorney for legal content, not just any available person); it checks for the specific known risk in that category (bias in a hiring rationale, a fabricated citation in a legal brief, a wrong dosage in medical text) rather than a general glance for typos; and it happens before the output is acted on or communicated, with a real ability to stop or change the outcome — not a review logged after the decision has already taken effect.
Key Concept
Hiring, medical, legal, and financial decisions affecting individuals are recurring high-stakes categories where full autonomy is inappropriate. Meaningful human review is performed by someone qualified for that specific content, checks for the category's specific known risk, and happens before the output is acted on — not a generic after-the-fact glance.
Common Exam Distractor
Watch for answers that reframe a high-stakes risk as a speed, formatting, or minor-accuracy concern ("the model might be slow," "it might misspell a name"). The substantive risk in these categories is almost always about biased or fabricated content affecting a real person's outcome — and the fix is human review, not better prompting alone, which reduces but doesn't eliminate the risk.