Study guides / CCAO-F / Domain 2

Output Evaluation and Validation · Lesson 1 of 7

2.1 — Detecting Hallucinations, Inconsistencies, and Bias

Spot plausible-sounding but wrong claims, internal contradictions, and one-sided framing in Claude's output — even without a gold-standard answer to check against.

A hallucination is a claim Claude states with the same confident, fluent tone as everything else in a response, even though it's fabricated, wrong, or unsupported. This is one of the harder evaluation skills to build, because the exam (and real use) rarely hands you a gold-standard answer key to check against. Most of the time, you're reading a single response and have to judge, from the response itself, whether something in it deserves a second look.

Fortunately, hallucinations, internal contradictions, and biased framing all leave detectable traces in the text itself — you don't need external verification to notice the pattern, only to confirm it.

Red Flags Without a Ground Truth

A handful of patterns are worth training yourself to notice on every pass:

Internal Contradictions and One-Sided Framing

Two more red flags don't require any outside knowledge at all — they're visible just from reading the response carefully, start to finish:

Internal contradiction: a response that states a total, a conclusion, or a fact early on, then states something incompatible with it later — a number in a summary that doesn't match a number in a table further down, or a line saying "there's no clear evidence for X" followed later by a paragraph that treats X as established. Reading the whole response in one pass, rather than skimming the conclusion, is what catches this.

Biased or one-sided framing: a response that presents a recommendation or comparison while only surfacing the case for one side — favorable language for the preferred option, neutral or absent language for the alternative, no mention of an obvious tradeoff or counterargument. This isn't about factual accuracy at all; a one-sided answer can be built entirely from true statements and still mislead by omission.

Key Concept

None of these checks require a reference answer. Read for internal consistency (does the response agree with itself?), calibration (does confidence match how knowable the claim actually is?), and balance (are both sides of a real tradeoff represented?). These three questions catch a meaningful share of problems before any external fact-check is even needed.

Common Exam Distractor

Don't equate a confident, fluent, well-organized response with a correct one — fluency is a property of the writing, not the facts. Also don't mistake "off-topic" or "poorly formatted" for the specific failure mode of hallucination; a hallucination is a fabricated or unsupported claim delivered in an otherwise perfectly on-topic, well-formatted response, which is exactly what makes it harder to catch.

Exam traps

Practice question

While reviewing a Claude-drafted market analysis, you notice the summary states '68.4% of surveyed competitors have adopted this pricing model' with no source given, and you can't recall this being in any document you provided. What's the most appropriate response?

  • A Flag the figure as unverified and confirm it against an actual source before it goes into the report Correct

    An oddly specific, unsourced number is a classic hallucination red flag — the appropriate response is to verify it independently before trusting or repeating it, not to accept or reject it on the spot.

  • B Accept the figure as accurate, since a number that precise wouldn't likely be invented

    Precision is not evidence — a model can generate a specific-sounding number in a plausible shape without it being grounded in any real source.

  • C Ask Claude to restate the figure with more confidence so it reads better in the final report

    Making an unverified claim sound more confident increases the risk of it being trusted and repeated without addressing whether it's actually true.

  • D Remove the number but keep the surrounding sentence and conclusion unchanged

    If the number was central to the claim, simply deleting it without checking whether the underlying conclusion still holds leaves the more fundamental accuracy question unresolved.

Build exercise: Build a Red-Flag Checklist by Testing It Live

Beginner · 20 minutes

You'll practice:

  1. In claude.ai, ask Claude a question likely to invite an oddly specific but hard-to-verify number — for example, 'What percentage of small businesses in the Midwest switched CRM providers last year, and what was the average cost of switching?' Read the response looking specifically for precise figures with no stated source.

    This surfaces the most common hallucination red flag in a live setting: a fluent, specific-sounding statistic you have no independent way to place.

    You should see: A response containing at least one specific percentage or dollar figure stated with no citation or source attribution.

    Hints
    1. Pick a narrow, somewhat obscure question — the more niche the topic, the more likely Claude has to generate a plausible-sounding rather than well-grounded figure.
    2. Underline or note every number in the response before judging anything else.
    3. Ask yourself for each number: could I say exactly where this came from? If not, it's a candidate to verify.
  2. Ask Claude a multi-part question where an early part of the answer could conflict with a later part — for example, ask it to both summarize a scenario's total budget up front and then itemize the budget line by line further down. Read the full response end-to-end, not just the summary.

    This builds the habit of reading for internal consistency rather than stopping at the first confident-sounding line, which is where contradictions hide.

    You should see: Either a consistent response (a good outcome worth noting), or a case where the stated total and the itemized figures don't quite add up.

    Hints
    1. Deliberately ask for a summary total near the top and a detailed breakdown later, so there's an actual internal number to cross-check.
    2. Do the arithmetic yourself on the itemized figures and compare it to the stated total.
    3. If everything is consistent on the first try, ask a follow-up that makes the scenario more complex to increase the chance of catching a slip.
  3. Ask Claude to recommend one option between two genuine alternatives (e.g. 'should our team use a monthly or annual subscription for this tool?') and read the response specifically for whether it acknowledges a real downside of its recommended option.

    This tests for one-sided framing directly — a common way biased output shows up even when every individual statement in it is true.

    You should see: A recommendation that either fairly represents a tradeoff on both sides, or one that discusses only the advantages of its preferred option.

    Hints
    1. Pick a decision that has a genuine, well-known tradeoff (e.g. upfront cost vs. flexibility) so you can judge the answer against something real.
    2. Check specifically whether the non-recommended option gets any fair treatment at all, or only the recommended one does.
    3. If the first answer is well-balanced, ask Claude to argue only for one side and compare how differently it reads.

Sources