Red Flags for Hallucination, Contradiction & Bias
None of these checks require a reference answer — read for internal consistency, calibration, and balance.
- Oddly specific, unverifiable numbers — precision is not evidence. If you can't say where a figure came from, treat it as unverified.
- Confident tone on an inherently uncertain question — a well-calibrated answer hedges on genuinely uncertain topics; flat confidence with no hedge is a signal to check.
- Invented specifics — a named case, paper title, page number, or quote can be generated in a plausible shape without being real.
- Internal contradiction — an early total/conclusion conflicts with a later statement (e.g. a summary number that doesn't match a table). Only caught by reading the full response, not just the conclusion.
- One-sided framing — favorable language for one option, silence on an obvious tradeoff. Can be built entirely from true statements and still mislead by omission.
Exam trap: fluent, confident, well-formatted writing is a property of style, not accuracy. Don't equate polish with correctness.
Fact-Checking: Cross-Source Verification & Citation Traceability
- Check a claim against at least two independent sources, not one — a single source can itself be the origin of an error.
- If sources disagree, the disagreement is the finding — say so, don't pick whichever source you saw first.
- Prioritize verification effort toward claims that are surprising, unusually precise, or high-consequence (a headline stat, a legal claim, a budget-driving number) over routine, low-stakes claims.
- A quoted passage is not verified just because it's in quotation marks — quotation marks show traceability was attempted, not confirmed.
- Real verification: open the cited source and confirm the exact text appears there, in the context the answer implies.
Exam trap: treating "the response includes a quote/citation" as sufficient evidence of accuracy on its own.
Why Self-Grading Fails — and the Fix
The problem is structural, not effort-based. The same reasoning, knowledge gaps, and assumptions that produced an output are what get reused to check it — a wrong assumption doesn't become visible just because the same model looks again.
| Doesn't Fix It | Why Not |
|---|---|
| More detailed grading rubric | Same source of error, same model, still applying it |
| "Act as a skeptical critic" in the same conversation | Persona changes tone, not the underlying reasoning or knowledge |
| Assuming it's only a risk at scale | A single self-graded response is just as susceptible as a batch of a thousand |
The actual fix — three ingredients, usable alone or combined:
- Independent, concrete criteria — a pass/fail checklist ("cites a source for each statistic?") instead of an open-ended "is this good?"
- A genuinely separate evaluation pass — a fresh conversation with explicit critic framing, not the same thread
- A human reviewer for anything consequential — the most reliable independent signal available
Reviewing Extended Thinking as a Validation Step
Extended thinking's value here is error-catching, not transparency for its own sake — a polished final answer can be confidently wrong in a way invisible from the answer alone, because the error happened steps earlier.
- Check the opening assumption first — everything built on it is compromised if it's wrong, no matter how sound later steps look.
- Verify numeric/comparative steps against numbers you actually provided, like checking a coworker's math.
- Confirm the final answer actually follows from the last reasoning step — watch for a subtle overreach at the end.
- Reserve this review for problems where it pays off: multi-step decisions, calculations with real consequences, multi-constraint comparisons. Skip it for simple lookups.
Exam trap: a longer or more detailed-looking trace is not automatically more reliable — length isn't a substitute for checking content.
Human Oversight for High-Stakes Decisions
| Category | Specific Known Risk |
|---|---|
| Hiring | Bias from input data correlated with protected characteristics, even unintentionally |
| Medical | Medically plausible-sounding text that's still wrong — the hallucination risk at its highest stakes |
| Legal | Fabricated case citations, statutes, or precedents — a well-documented failure mode |
| Financial (individuals) | Biased outcomes plus an accountability requirement to explain/stand behind an adverse decision |
What "additional verification" actually means — all three must be true, not just one:
- Performed by someone qualified for that specific content (a clinician, an attorney — not just any available person)
- Checks for the category's specific known risk, not a general glance for typos
- Happens before the output is acted on, with real ability to stop or change the outcome
Exam trap: reframing a high-stakes risk as a speed or formatting concern, or treating careful prompting alone as a substitute for human review — prompting reduces risk, it doesn't eliminate it.
Checking Semantic Consistency Without a Gold Answer
Ask the same underlying question two or three genuinely different ways and compare the substance of the answers, not the wording.
- Different wording, same substance = good — expected from a reliable answer expressed two ways.
- Same substance check: does the bottom-line conclusion, number, named recommendation, or a caveat match across phrasings?
- A substantive difference is a real finding — a different conclusion, number, or a caveat present in one answer and silently missing in another — worth independent verification.
- Avoid leading phrasings — if one version nudges toward an answer while the other stays neutral, a difference may reflect the leading wording, not genuine inconsistency.
- Faster technique: paste both answers into one conversation and ask Claude to identify substantive differences — narrower and more checkable than open-ended self-grading, but still verify it yourself.
Exam trap: judging consistency by word overlap or phrasing similarity instead of substance; treating the more detailed or later answer as automatically correct when two answers disagree.
Editing, Comparing Variants, and Choosing the Right Output Format
Treat a first draft as a starting point, not a finished deliverable.
- Revise against a named, specific audience gap — "this needs the recommendation up front, then supporting detail" beats "make it more concise."
- Generate genuinely distinct variants and compare them against the actual need — ask for what should differ (length, framing, detail level), not near-identical rewordings.
- Compare deliberately — against the audience's actual needs, not whichever variant you read first or sounds most polished.
| Format | Best For |
|---|---|
| Inline chat reply | A short, conversational answer read once, in context |
| Artifact | Standalone content the reader will reference, reuse, or edit outside the conversation (a document, report, diagram) |
| Structured data (table/JSON/CSV) | Information a person or tool needs to sort, filter, or compare across items along shared fields |
Exam trap: a vague "make it better"/"more professional" instruction, or defaulting to the same output format regardless of what the reader actually needs to do with it.