Study guides / CCAO-F

Glossary

Quick-lookup definitions for every domain, with exam context and links back to the lesson that covers each term.

Hallucination
A claim Claude states with the same confident, fluent tone as the rest of a response, even though it is fabricated, wrong, or unsupported. Common shapes include oddly specific unverifiable numbers, confident phrasing on an inherently uncertain question, and invented specifics (a named case, a paper title, a quote) that sound plausible without corresponding to anything real.
Exam context: The exam tests whether you can spot hallucination without a gold-standard answer key, by reading for red flags rather than checking against a reference. A common distractor equates a confident, well-organized response with a correct one, or confuses hallucination with an off-topic/poorly-formatted answer.
See also: 2.1 Detecting Hallucinations, Inconsistencies, and Bias
Internal Contradiction
A response that states a total, conclusion, or fact early on, then states something incompatible with it later in the same response — for example, a summary number that doesn't match a table further down. Catching this requires reading the whole response start to finish rather than skimming only the conclusion.
Exam context: Expect a scenario where an early statement conflicts with a later one; the correct behavior is reading the full response for consistency, not stopping at the final line.
See also: 2.1 Detecting Hallucinations, Inconsistencies, and Bias
One-Sided (Biased) Framing
A response that presents a recommendation or comparison while surfacing the case for only one side — favorable language for the preferred option, little or no mention of an obvious tradeoff or counterargument. This is not about factual accuracy; a one-sided answer can be built entirely from true statements and still mislead by omission.
Exam context: A common distractor is treating a one-sided recommendation as automatically reliable because every individual statement in it is true. The exam wants you to catch selective emphasis independent of accuracy.
See also: 2.1 Detecting Hallucinations, Inconsistencies, and Bias
Cross-Source Verification
Checking a claim against at least two independent sources before treating it as settled, rather than relying on a single source that could itself be the origin of an error. When sources disagree, the disagreement is itself the finding — the honest answer is that sources disagree, not a confident pick of whichever source was seen first.
Exam context: Expect scenarios that reward prioritizing verification effort toward surprising, precise, or high-consequence claims rather than checking every claim with equal rigor. A single-source check is a common wrong answer.
See also: 2.2 Fact-Checking and Citation-Traceable Answers
Citation-Traceable Answer
An answer where a claim is backed by a quoted passage that has actually been verified by opening the cited source and confirming the exact wording appears there, in context. Quotation marks around a passage make it look traceable, but only opening the source and checking makes it verified.
Exam context: A frequent distractor treats "the response includes a quote" or cites a source as sufficient evidence of accuracy on its own. The exam expects you to name the actual verification step: opening the source and checking.
See also: 2.2 Fact-Checking and Citation-Traceable Answers
Self-Grading Bias
The unreliability that results when the same model instance that generated an output is also asked to evaluate it. The problem is structural, not a matter of effort: the same reasoning patterns, knowledge gaps, and assumptions that produced the original output are reused when grading it, so a wrong assumption made the first time is unlikely to be caught the second time by the same reasoning.
Exam context: A classic distractor is "fixing" self-grading with a more detailed rubric or a same-conversation 'act as a skeptical critic' instruction — both still run through the same model instance and don't remove the structural cause. The exam expects you to recognize that only a change in who or what does the grading solves this.
See also: 2.3 The Risk of a Model Grading Its Own Output
Independent Evaluation Pattern
A more reliable alternative to self-grading with three ingredients usable alone or in combination: independent, concrete pass/fail criteria instead of an open-ended "is this good?" judgment; a genuinely separate evaluation pass (a fresh conversation with explicit critic framing); and a human reviewer for anything consequential.
Exam context: Expect a scenario asking you to design an evaluation approach; the correct answer combines concrete criteria with a fresh, separate pass and, for high-stakes output, a human check — not just a stricter prompt to the same conversation.
See also: 2.3 The Risk of a Model Grading Its Own Output
Extended Thinking (as a Validation Step)
Claude's visible, step-by-step reasoning shown before a final answer when extended thinking is enabled. As a validation technique, its value is catching a flawed opening assumption, a bad numeric or comparative step, or a final answer that doesn't actually follow from the reasoning shown — before the polished final answer gets trusted and acted on.
Exam context: A common distractor treats a longer or more detailed trace as automatically more trustworthy, or frames reviewing the trace as being about speed or making an answer merely feel more trustworthy. The exam rewards reserving this review for complex, multi-step, or high-stakes problems, not every task.
See also: 2.4 Reviewing Extended Thinking as a Validation Step
Human Oversight for High-Stakes Decisions
The practice of requiring meaningful human review before Claude output is acted on in categories where consequences to a real person are serious: hiring, medical, legal, and financial decisions affecting individuals. Meaningful review is performed by someone qualified to judge that specific content, checks for the category's specific known risk (bias, fabricated citation, wrong dosage, etc.), and happens before the output is acted on, with real ability to stop or change the outcome.
Exam context: Expect the exam to reframe the risk as a minor speed or formatting concern as a wrong answer, and to treat better prompting alone, or a generic after-the-fact glance, as an insufficient substitute for a qualified reviewer checking before action is taken.
See also: 2.5 Human Oversight for High-Stakes Decisions
Semantic Consistency (Across Paraphrases)
A reliability check where the same underlying question is asked two or three different ways and the answers are compared for agreement in substance — the bottom-line conclusion, number, or recommendation — not for similarity of wording. Agreement despite different wording is a good sign; a different substantive conclusion across phrasings is a reliability red flag worth verifying independently.
Exam context: A common distractor judges consistency by word overlap or phrasing similarity rather than substance, or treats a more detailed or later answer as automatically correct when two paraphrased answers disagree. Watch also for leading phrasings that bias one version toward a particular answer.
See also: 2.6 Checking Semantic Consistency Across Paraphrases
Audience-Targeted Revision
Revising a Claude draft against a concrete, named gap in what a specific reader knows and needs, rather than a vague instruction like "make it better." Naming the actual gap (e.g. "this needs the recommendation up front, then supporting detail after") produces a targeted edit; a vague instruction tends to produce only a differently-worded draft.
Exam context: Expect scenarios contrasting a vague revision request with an audience-specific one; the exam rewards naming what a specific reader needs to walk away knowing or deciding.
See also: 2.7 Editing, Refining, and Comparing Outputs
Output Format Selection
Choosing among an inline chat reply, an Artifact, or structured data (a table or JSON/CSV-style export) based on how the reader will actually use the output — read once in context, referenced or reused standalone, or sorted and compared across items — rather than defaulting to whatever format was used last time.
Exam context: The exam treats format choice as part of curating output for a reader: the right content in the wrong format still fails the reader. Defaulting to the same format regardless of use case is a common wrong answer.
See also: 2.7 Editing, Refining, and Comparing Outputs