Skip to content

Why LLMs hallucinate and what actually reduces it

By SunnyKumar Jonwal 9 min read

A language model states a court case that doesn't exist, quotes a paragraph from a paper that was never written, or calls a method your library doesn't have, and does it in the same confident tone it uses for true statements. That's a hallucination, and it's the most cited reason people distrust these systems.

It isn't a bug that will be patched out next month, and it isn't random either. It follows from what these models are and how they're trained. Once you see the mechanism, the fixes make sense, and so do the limits of each.

What's actually happening

A language model is trained to predict the next piece of text given what came before. Over an enormous amount of text, that training produces something that knows a great deal about language, facts, code, and reasoning patterns. But the objective is producing plausible continuations, and plausibility isn't the same as truth.

When the model knows the answer well, the most plausible continuation is also the correct one. When it doesn't, for a rare fact or a niche API or a recent event, it still produces a fluent, plausible-sounding continuation, because that's what it does. There's no built-in "I don't have this" alarm that reliably fires. The confidence in the tone comes from the fluency, and it isn't a measure of how much evidence backs the claim.

Two more forces push in the same direction. Models trained to be helpful learn that answering is usually rewarded and refusing usually isn't. And some researchers have argued that common ways of evaluating models, which score an answer as right or wrong but give no credit for saying "I'm not sure," effectively reward guessing over abstaining. A system tuned against tests like that learns to guess.

The main types

It's worth separating the kinds, because they call for different fixes.

Fabricated facts. Invented dates, statistics, biographies, quotes, or events. Most common for rare or obscure topics, where the training data was thin.

Fabricated sources. Made-up citations, URLs, paper titles, and case law that look real. Lawyers have been sanctioned for filing briefs containing invented case citations generated by a chatbot, which turned an abstract concern into a courtroom problem.

Invented code and APIs. Functions, parameters, and packages that don't exist, or that exist in a different version. Plausible naming conventions make these easy to generate and easy to overlook, and as covered in AI coding risks, invented package names can become a security issue too.

Reasoning errors. Steps that look logical and aren't, arithmetic slips, or conclusions that don't follow. The prose reads as sound reasoning while the substance is broken.

Context errors. Ignoring or distorting the information you supplied. A model asked for a summary of a document may add details that aren't in it, or contradict what it says.

Sycophancy. Agreeing with a false premise in the question, or with the user's stated opinion, because agreement is a likely continuation. "Why did Einstein win the Nobel for relativity?" invites an answer built on the wrong premise.

Why it's hard to eliminate

A few structural reasons this persists.

The model has no direct access to ground truth. It has patterns learned from text, and those patterns include errors and disagreements. It can't check its answer against the world unless you connect it to something that can.

The long tail is endless. Popular facts appear thousands of times in training data and are learned well. Facts that appear once, or never, are where errors cluster, and any real application eventually asks about them.

Fluency hides the error. A wrong answer written in perfect prose looks exactly like a right one. Humans read confidence as competence, which is a poor heuristic for machines.

And no setting removes it. Turning temperature to zero makes output more deterministic, and a deterministic wrong answer is still wrong. Bigger, newer models tend to hallucinate less often on many measures, but none has stopped.

What helps: ground the model in real information

The single most effective technique is to stop asking the model to remember things and give it the facts instead. Put the relevant source text in the prompt and instruct the model to answer only from it. That's retrieval-augmented generation, described in RAG that works, and it works because the model is now doing something it's good at, reading and summarizing, instead of something it's bad at, recalling obscure details from memory.

Grounding isn't a cure. Retrieval can fetch the wrong text, and the model can still misread or embellish. But it shifts the problem from "is the model's memory accurate?" to "is our retrieval good and is the answer supported?", and those you can measure and fix.

Give it tools for the things it's bad at

Models are unreliable at arithmetic, current information, precise counting, and exact lookups. Don't fight that. Give them a calculator, a search tool, a database query, a code interpreter. When the model decides to call a tool and reads the result, it stops guessing. The mechanics are in tool use explained.

The prompt should say when tools are mandatory: "Never state a price from memory. Always call get_price." Otherwise the model may answer from memory when it feels confident, which is the situation you were worried about.

Make it show its work

A few prompt-level habits reduce fabrication noticeably.

Ask for quotes before conclusions. "First extract the exact sentences from the documents that answer the question, then answer using only those." Now every claim has to trace back to text, and you can check programmatically that the quoted text appears in the source.

Ask for citations by ID, then verify that each cited source actually supports the claim, at least on a sample.

Give it an exit. "If the documents don't contain the answer, say 'not found.'" and mean it: test that it uses the exit. Without explicit permission to abstain, the model tends to fill gaps.

Ask it to flag uncertainty. Models are imperfect at self-assessment, but prompting for "what would you need to verify this?" or "which parts are you least sure about?" often surfaces weak spots. Treat the response as a hint, not a guarantee.

More prompt-level guidance is in prompt engineering that still works.

Constrain the output

The less freedom the model has, the fewer places to invent. If the answer must be one of five categories, use an enum. If a field must be a date, validate it. If it must reference an existing record, check the ID against your database, and reject any that don't exist. Schema-constrained output, covered in getting reliable JSON out of an LLM, turns some hallucinations into validation errors you can handle.

For code, let real tools judge the output: compile it, run the tests, run the linter, resolve the imports. A hallucinated function fails immediately when executed, which is a much better failure than a reviewer missing it.

Verify with a second pass

You can use one model call to check another. Common patterns:

  • Self-consistency. Ask the same question several times and compare. Facts that vary across samples are suspect, since a model that knows something tends to say it consistently while a guess wanders.
  • A critic pass. A second call receives the answer and the source and is asked to list claims not supported by the source.
  • Cross-checking with retrieval. For each claim in an answer, search for support and flag the ones with none.

These help, with caveats. A checking pass can share the same blind spots as the first, and it adds cost and latency. It works best when the checker has something the generator didn't, like the source text or a tool.

Design the product around fallibility

Since you can't reduce the rate to zero, design for the leftover.

Show sources next to claims, so users can check. Make it easy to see and follow citations. Present confidence honestly, and avoid the tone of an oracle. Keep humans in the loop for high-stakes outputs: medical, legal, financial, security, anything where a wrong answer causes real harm. Let users report errors, and feed those reports back into your test set. And be candid in the interface that the system can be wrong, in language people will actually notice.

Match the stakes to the design. A brainstorming assistant can tolerate an occasional invention. A tool summarizing contracts can't.

Measure it

You can't manage what you don't measure, and hallucination is measurable on your own data. Build a set of questions where you know the answers, including some the system should decline. Include cases with false premises, questions about things not in the documents, and requests for citations. Then score: is the answer correct, is every claim supported by the provided sources, did it abstain when it should have?

Track the rate over time and across changes to the prompt, model, retrieval, or tools. A change that improves fluency while quietly raising the fabrication rate is easy to ship if you're not counting. The approach in how to evaluate an AI agent applies here, including calibrated LLM judges for the open-ended cases and human review of samples.

A worked example: a support bot that invents refund rules

Picture a support assistant for an online shop. A customer asks whether a discounted item can be returned after 45 days. The policy document says returns are accepted within 30 days, with an exception for defective goods. The bot answers, warmly and confidently, that discounted items can be returned within 60 days. Nothing in the documents says that. The number is a plausible blend of common return windows the model has seen across the internet.

Trace the failure and each fix becomes obvious. The model answered from memory because nothing forced it to use the policy text, so the first change is to retrieve the relevant policy section and put it in the prompt. It still might embellish, so the second change is to require a quoted sentence from the policy alongside every answer. It might quote nothing when the policy is silent, so the third change is an explicit exit: if no passage answers the question, hand the conversation to a person. Finally, a test set of fifty real questions, ten of which the policy doesn't cover, tells you whether the fixes worked or just moved the problem.

None of those steps needed a better model. They needed a better system around the model, and that's usually where the gains are.

Myths worth dropping

"Lower the temperature and it stops." It makes output more repeatable, and it doesn't make it more accurate.

"The newest model doesn't hallucinate." It hallucinates less on many tasks, and it still does, especially on obscure or recent material.

"RAG solves it." RAG reduces it substantially when retrieval is good, and it introduces its own failure modes.

"It only happens on hard questions." It also happens on easy-looking ones, especially those with a false premise or a request for a specific citation.

"If it sounds sure, it's right." Tone tells you nothing about accuracy.

The practical stance

Treat a language model as a brilliant, fast, occasionally overconfident colleague who has read a lot and can't check anything on their own. You wouldn't let that person sign off on a legal filing without verification, and you would happily let them draft it. Give them sources to work from and tools to check things with, ask them to show their evidence, and review what matters.

Do those things and hallucinations become a manageable engineering problem, with measurable rates, known causes, and mitigations that each shave off a piece. Ignore them and they show up as the embarrassing screenshot.