A system prompt is a contract with the model, not a wish list. It has three jobs: give the model the role and context it lacks, define the decision criteria it must apply, and specify the output it must produce. It also has a hard limit: it shapes behaviour probabilistically and cannot enforce anything. Most scenario questions in this area hinge on separating those two facts.
Explicit criteria beat vague dials
Instructions such as “be conservative” or “only report high-confidence findings” sound rigorous but give the model no decision boundary. Anthropic’s prompting guidance says the same thing from the other direction: be clear and direct, explain why a rule exists (the model generalises from the reason), and tell it what to do rather than only what to avoid. Its “golden rule” is a good review test: show the prompt to a colleague with no context; if they would be confused, the model will be too.
For a review or triage task, replace dials with categorical criteria and calibrate them with concrete examples rather than prose descriptions:
Report: bugs, security vulnerabilities, and comments whose
claimed behaviour contradicts what the code does.
Skip: style preferences and naming that matches the local pattern.
Severity examples:
critical - user input concatenated into a SQL string
major - property access on a value that can be null
minor - inconsistent variable naming inside one moduleThere is a model-behaviour twist worth knowing. Anthropic notes that on Claude Opus 5 an instruction like “only report high-severity issues” or “be conservative” may be followed literally and simply report less; its recommendation is to ask for everything (with a severity or category field) and filter in a separate pass. That gives you a measurable filter you control. Self-reported confidence is a weak substitute: use it to route items to review only after you have checked it against labelled data.
Trust is system-wide, so control quality per category
When one output category has a high false-positive rate, users stop trusting every category, including the accurate ones. The counter-intuitive fix is to switch the noisy category off while you rework its criteria and examples, then re-enable it once measured precision recovers. Track precision per category, not one blended number, so the noisy category is visible (evaluation design is covered in 3.1 and 3.2).
Structure and templating
Treat a production prompt as a template with a stable skeleton and clearly delimited slots for variable content.
- Separate instructions from data. Wrap each kind of content in its own tag (for example
<instructions>,<context>,<input>) with consistent, descriptive names. The docs use double-brace placeholders such as{{CONTENT}}in their own templates; the syntax is a convention, the separation is the point. - Order for two audiences. Put long documents above the query and instructions (see 6.4). Put stable content first and per-request values last, so the same layout supports prompt caching (see 6.5). A timestamp or user name inside the stable region silently defeats the cache.
- Match style to output. Prompt formatting influences response formatting; describe the format you want positively (“smoothly flowing prose”) rather than banning what you do not want.
- Dial back shouting. Newer models respond strongly to the system prompt, and emphatic language written to overcome undertriggering in older models can now cause overtriggering. Use normal wording.
- Do not rely on prefill. Prefilled responses on the last assistant turn are not supported starting with Claude 4.6 models; use direct instructions or structured outputs.
Guardrail wording versus enforcement
Key concept: probabilistic versus deterministic controls
Wording in a system prompt is a probabilistic control: it lowers the likelihood of bad behaviour and shapes how the model refuses. Permissions, schemas, validators, sandboxing and least-privilege credentials are deterministic controls. Any requirement of the form “this must never happen” (a refund above a limit, a destructive tool call without approval) belongs in the second category, with the prompt as an additional layer, not the only one. See 4.1 for the safety-control view and 1.2 for authorisation.
Anthropic’s guidance on jailbreaks and prompt injection is layered. For hostile users: pre-screen input with a lightweight model (the docs use Claude Haiku 4.5) constrained to a simple classification via structured outputs, validate input, write a system prompt that states boundaries and exactly how to refuse, and monitor and respond to repeat offenders. For hostile content (indirect injection through web pages, emails or tool results): deliver third-party content only inside tool_result blocks, say what it is and where it came from, state in the system prompt that tool and document content is untrusted data that must never override instructions, JSON-encode untrusted strings, screen tool output before the model acts on it, apply least privilege, and red-team the agent before deployment.
Structured outputs (output_config.format with a JSON schema) guarantee schema-valid output through constrained decoding, but they do not guarantee that a value is true, the model can still refuse, and a low max_tokens can truncate the JSON. Schema validity is a format control, not a correctness control.
Version prompts like code
Keep prompts in source control, review changes by pull request, and record with each version the model ID and eval results it was validated against. This matters because prompts are coupled to models: Anthropic’s guidance treats model-specific advice as measured on that model and says to re-check it against your own evals before applying it elsewhere. Concrete examples from the docs: explicit “verify your work” instructions can cause over-verification on Opus 5, and aggressive tool-use language tuned for older models can overtrigger on newer ones. Log the prompt version with each request (see 1.4), roll changes out with an A/B comparison (3.3), and keep the previous version deployable for rollback.
Common exam distractor
“Add a line to the system prompt telling the model to be conservative, never do X, or only output high-confidence results” is the tempting wrong answer to almost every precision, safety and false-positive question. The better answer pairs explicit categorical criteria (and examples) for behaviour with deterministic enforcement in code for anything that must be guaranteed.