A hallucination is a claim Claude states with the same confident, fluent tone as everything else in a response, even though it's fabricated, wrong, or unsupported. This is one of the harder evaluation skills to build, because the exam (and real use) rarely hands you a gold-standard answer key to check against. Most of the time, you're reading a single response and have to judge, from the response itself, whether something in it deserves a second look.
Fortunately, hallucinations, internal contradictions, and biased framing all leave detectable traces in the text itself — you don't need external verification to notice the pattern, only to confirm it.
Red Flags Without a Ground Truth
A handful of patterns are worth training yourself to notice on every pass:
- Oddly specific, unverifiable numbers. A precise-sounding figure — "62.3% of mid-sized firms," "the policy took effect in March 2019, section 4.2(b)" — reads as more credible than a round one, but precision is not evidence. If you can't immediately place where that exact figure would come from, treat it as a claim to verify, not a fact to repeat.
- Confident tone on an inherently uncertain question. A well-calibrated answer hedges when the underlying question is genuinely uncertain or outside what's knowable from training data (a very recent event, a niche internal policy, a prediction). An answer that states an inherently uncertain thing flatly, with no hedge at all, is a signal worth noticing — not proof of an error, but a reason to check.
- Invented specifics that sound plausible. A named case, a paper title, a page number, a quote attributed to a real person — these are exactly the kind of detail a model can generate in a plausible-sounding shape without it corresponding to anything real. Plausible shape is not the same as real.
Internal Contradictions and One-Sided Framing
Two more red flags don't require any outside knowledge at all — they're visible just from reading the response carefully, start to finish:
Internal contradiction: a response that states a total, a conclusion, or a fact early on, then states something incompatible with it later — a number in a summary that doesn't match a number in a table further down, or a line saying "there's no clear evidence for X" followed later by a paragraph that treats X as established. Reading the whole response in one pass, rather than skimming the conclusion, is what catches this.
Biased or one-sided framing: a response that presents a recommendation or comparison while only surfacing the case for one side — favorable language for the preferred option, neutral or absent language for the alternative, no mention of an obvious tradeoff or counterargument. This isn't about factual accuracy at all; a one-sided answer can be built entirely from true statements and still mislead by omission.
Key Concept
None of these checks require a reference answer. Read for internal consistency (does the response agree with itself?), calibration (does confidence match how knowable the claim actually is?), and balance (are both sides of a real tradeoff represented?). These three questions catch a meaningful share of problems before any external fact-check is even needed.
Common Exam Distractor
Don't equate a confident, fluent, well-organized response with a correct one — fluency is a property of the writing, not the facts. Also don't mistake "off-topic" or "poorly formatted" for the specific failure mode of hallucination; a hallucination is a fabricated or unsupported claim delivered in an otherwise perfectly on-topic, well-formatted response, which is exactly what makes it harder to catch.