New: The 5-Day Stoic Operator Challenge — Free. Start today →

AI Hallucinations: Why Models Make Things Up and the Operator's Verification Protocol

AI Hallucinations: Why Models Make Things Up and the Operator's Verification Protocol

An AI hallucination is not a glitch you wait out. It is a predictable output of how language models are built and graded, and an operator manages it the way a pilot manages crosswind: by design, every time.

You already use AI to draft proposals, summarise calls, write SOPs and answer client questions. Most of that output is correct. The problem is the small share that is wrong and reads exactly like the share that is right.

A wrong answer that sounds unsure costs you nothing, because you check it. A wrong answer that sounds certain costs you a client, a refund or a court filing. That is the whole risk in one line, and it is why AI hallucinations deserve a protocol rather than a shrug.

This article covers what a hallucination is according to the people who measure them, why models produce them, the real damage they have already done, and the verification protocol an operator runs before AI output reaches a decision.

What AI hallucinations are

The working definition on Wikipedia's entry on AI hallucination describes a response that contains "false or misleading information presented as fact". The key word is presented. The model does not flag the invented part. It delivers fact and fiction in the same calm register.

The US National Institute of Standards and Technology prefers the word confabulation. Its Generative AI Profile (NIST AI 600-1) lists it as a core model risk and defines it as a phenomenon in which systems "generate and confidently present erroneous or false content in response to prompts". The same document widens the net to outputs "that diverge from the prompts or other input or that contradict previously generated statements in the same context".

That second clause matters for operators. A hallucination is not only an invented fact about the world. It is also a summary that drifts from the document you supplied, or a reply that contradicts what the model told you three messages earlier.

Three shapes to recognise

  • Invented facts. A statistic, a date, a price or a feature that does not exist.
  • Invented sources. Citations, case law, studies or URLs that look real and lead nowhere.
  • Unfaithful summaries. Output that misstates or adds to the material you gave it.

Why language models hallucinate

The cause is structural, not mysterious. NIST states it plainly: "Confabulations are a natural result of the way generative models are designed". The models produce text that approximates the statistical patterns of their training data. Most of the time the likely answer is the true one. Sometimes the likely-sounding answer is false.

A September 2025 paper, Why Language Models Hallucinate, adds the second half of the explanation. It compares models to students on a hard exam who guess, "producing plausible yet incorrect statements instead of admitting uncertainty". The authors argue the habit survives because of how models are scored: "language models are optimized to be good test-takers, and guessing when uncertain improves test performance".

Read that as an incentive problem. If a benchmark gives zero points for "I don't know" and some points for a lucky guess, a system tuned to win the benchmark learns to guess. The paper's proposed fix is to change how existing benchmarks are scored, not to bolt on more hallucination tests.

The operator lesson is first principles. You cannot fix the incentive inside a vendor's training pipeline. You can change the incentive inside your own prompt, by making "I don't know" an acceptable and rewarded answer. More on that in the protocol.

Why confidence is not evidence

NIST flags a second trap: outputs "may also include confabulated logic or citations that purport to justify or explain the system's answer". A neat chain of reasoning and a tidy reference list are not proof. They can be generated by the same process that generated the error.

This is the Stoic discipline of assent applied to software. An impression arrives looking true. The operator examines it before agreeing to it. The model's tone is an impression, not a verdict.

What hallucinations have already cost

The damage is documented, and it lands on the person who used the output, not the model.

  • A court filing. Per Wikipedia's account, in 2023 a lawyer in a personal injury case against the airline Avianca submitted "six fake case precedents generated by ChatGPT" to a federal court in New York. The judge dismissed the case and fined two lawyers $5,000 for bad faith conduct.
  • A customer policy. The same entry records that in February 2024 a Canadian tribunal ordered Air Canada to pay damages and honour a bereavement fare policy "that was hallucinated by a support chatbot". The tribunal rejected the airline's argument that the chatbot was a separate legal entity responsible for its own actions.

Note the pattern in both. The person or company that shipped the output owned the consequence. A support bot on your site speaks for your business. A proposal drafted by AI and sent under your name is your proposal.

The research community sees the same exposure. The 2024 Nature paper on semantic entropy from a University of Oxford team lists problems including "fabrication of legal precedents" and untrue facts in news articles, and notes that hallucination can even pose a risk to human life in medical domains such as radiology.

How researchers detect hallucinations

The Oxford team targeted a specific subset of hallucinations they call confabulations, defined as "arbitrary and incorrect generations". Their method asks the model the same question several times and measures how much the answers disagree, computing "uncertainty at the level of meaning rather than specific sequences of words".

The intuition transfers directly to your desk. When a model knows something, repeated answers converge on the same meaning even if the wording varies. When it is guessing, the answers scatter. Scatter is a warning light.

Anthropic's own guidance on reducing hallucinations recommends the manual version of this, which it calls best-of-N verification: run the same prompt multiple times and compare, because inconsistencies across outputs could indicate hallucinations. You do not need a research lab to apply it. You need three tabs and two minutes.

The operator's hallucination protocol

The protocol below is built from the techniques Anthropic documents as of October 2026, ordered by when you apply them. It is written for Claude because that is the stack this desk uses, but the moves apply to any model. For the wider prompting toolkit, see the prompt engineering guide for operators.

Step 1: Tier the task before you prompt

  1. Tier 1, low stakes. Brainstorms, first drafts, subject lines, outlines. Errors are cheap and visible. Light review.
  2. Tier 2, internal decisions. Summaries of calls, data interpretation, SOP drafts. Grounding plus a spot-check of every number.
  3. Tier 3, external or irreversible. Anything sent to a client, published, filed, priced or touching health, money or law. Full protocol, plus a human check of every claim against a source.

Most operators skip this step, then apply Tier 1 review to Tier 3 output. That is the failure mode in both documented cases above.

Step 2: Give permission to not know

Anthropic's first basic technique is to explicitly give the model permission to admit uncertainty, which the documentation says "can drastically reduce false information". This is the direct counter to the test-taker incentive. Add a line like this to every Tier 2 and Tier 3 prompt:

  • If the material does not contain the answer, or you are not confident, say 'I don't have enough information' and stop. A gap is acceptable. A guess is not.

Step 3: Ground the model in your documents

Two of Anthropic's techniques work together here. First, restrict the model to the material you supply rather than its general knowledge. Second, for long documents (the guide gives the threshold as over 20k tokens), ask it to extract word-for-word quotes first and then reason only from those quotes.

  • Using only the attached contract, first list the exact quotes relevant to termination terms, numbered. Then answer my question using only those quotes, citing them by number. If no relevant quote exists, write 'No relevant quotes found.'

If you build on the API, use the native Citations feature. Anthropic's Citations documentation states that because the API extracts the cited text directly, "citations are guaranteed to contain valid pointers to the provided documents". That guarantee covers the pointer, not your judgement. You still read the cited passage.

Step 4: Make the model audit itself

After the draft, run a second pass. Anthropic's guide describes having the model find a supporting quote for each claim after it generates a response: "If it can't find a quote, it must retract the claim." In practice:

  • Review each factual claim in your draft. For each one, find the direct quote from the documents that supports it. Remove any claim without a supporting quote and mark the spot with [].

Empty brackets are a gift. They show you exactly where the model was about to make something up.

Step 5: Check for scatter

For Tier 3 questions where the answer comes from the model's own knowledge, ask the same question in three fresh sessions. If the three answers agree in meaning, confidence rises. If they disagree on a number, a name or a date, treat every version as unverified until you find a primary source.

Step 6: Verify the load-bearing claims yourself

Anthropic's guide closes with the line that should sit above every operator's desk: "while these techniques significantly reduce hallucinations, they don't eliminate them entirely". Its instruction is to always validate critical information, especially for high-stakes decisions.

  1. List the claims a decision rests on. Usually three to five.
  2. Open every cited URL. Confirm it loads and confirm the quoted words appear on the page.
  3. Check every number against the original source, not against the model's restatement of it.
  4. Anything you cannot verify is cut or labelled "not established". Never filled with a plausible guess.

The same rule applies when you delegate to agents rather than chat. This desk wrote up that failure mode in detail in what your AI agent made up.

The Stoic frame: control what is yours

The dichotomy of control sorts this cleanly. Whether a model hallucinates is not up to you. Which output you accept, sign and send is entirely up to you. The operator stops arguing with the first category and builds a system around the second.

Premeditatio malorum does the rest. Before a Tier 3 deliverable ships, ask what the worst plausible error in it would be and where it would surface. The Stoic risk-planning method takes ten minutes and replaces the vague hope that the model got it right with a specific check.

This is the same discipline that runs your training. You do not trust a feeling that a set was heavy. You log the weight. You do not trust a feeling that an answer is right. You check the source.

Frequently asked questions

What is an AI hallucination?

An AI hallucination is output from a language model that contains false or misleading information delivered as if it were fact. It includes invented statistics, fake citations and summaries that misstate the source document. NIST calls the same phenomenon confabulation and lists it as a core risk of generative AI. The defining feature is confidence: the false part reads exactly like the true part.

Why do AI models hallucinate?

Models generate text that approximates patterns in their training data, so a likely-sounding answer can be false. A 2025 paper argues that training and evaluation also reward guessing over admitting uncertainty, because most benchmarks score a lucky guess higher than an honest "I don't know". The result is a system that fills gaps with plausible content instead of flagging them.

Can you stop AI hallucinations completely?

No. Anthropic's own documentation states that its recommended techniques significantly reduce hallucinations but do not eliminate them entirely, and it advises validating critical information, especially for high-stakes decisions. Grounding the model in your documents, allowing it to say it does not know and requiring quotes for each claim lowers the rate. Human verification of load-bearing claims remains the final control.

How do I check if an AI answer is a hallucination?

Open every source the model cites and confirm the quoted words actually appear there. Check each number against the original document, not the model's summary. Ask the same question in three fresh sessions and compare: answers that disagree on a name, date or figure signal guessing. Anything you cannot confirm should be cut or marked as not established.

Which prompts reduce hallucinations the most?

Three moves from Anthropic's guidance carry most of the weight. Give the model explicit permission to say it lacks enough information. Restrict it to documents you provide and have it extract exact quotes before answering. Then ask it to find a supporting quote for each claim in its draft and retract any claim it cannot support. Together these turn silent guesses into visible gaps.

Who is responsible when an AI chatbot gives wrong information?

In the documented cases, the business or professional who used the output. A Canadian tribunal ordered Air Canada to honour a policy its support chatbot invented and rejected the argument that the chatbot was a separate legal entity. In a 2023 US federal case, lawyers were fined after filing AI-generated fake precedents. Treat AI output sent under your name as your own statement.

Leverage with a verification layer

AI is leverage for the operator, and leverage multiplies whatever it is applied to, including errors. The model will keep producing confident falsehoods at some rate. The question is whether your system catches them before a client, a court or a customer does.

Tier the task, permit "I don't know", ground the model, make it audit itself, check for scatter and verify what the decision rests on. Run that until it is reflex, the way a warm-up is reflex before a heavy set. For more on putting the tool to work without letting it run you, read why the tool is not the bottleneck.

The discipline that makes AI safe to use is the same discipline that builds the body and steadies the mind. If you want to install it as a daily system, start with the free 5-Day Stoic Operator Challenge.

AIai hallucinationsai leverageapex life fitnesscompound performanceprompt engineeringverification
TH

The Apex Desk

The editorial team behind Apex Life Fitness — operators writing about the systems where fitness, philosophy, and AI leverage intersect. Train. Think. Build.