The dangerous AI answer is not the one that looks wrong. An obviously wrong answer gets caught in seconds. The dangerous one is fluent, confident and plausible, and it survives a quick read because everything about its surface says "correct". In tax research, where a single misread threshold or a case cited for the wrong point can flow straight into a client file, that surface plausibility is exactly the problem.
These fluent errors have a name: hallucinations. They are not a sign that a model is broken. They are a predictable feature of how general-purpose language models work, and they matter most precisely where the stakes are highest. Understanding why they happen, and building the habits and tools to catch them, is now a core professional skill.
This guide explains what a hallucination is, why it happens, the three habits that catch most of them, and how to cut the risk at source with tools built for verification.
The word covers a range of failures that share one feature: the output is confident and the content is false.
The most notorious form is an invented citation: a case, a section number or a piece of guidance that does not exist, or that exists but says something entirely different. Legal and tax contexts have already seen real-world consequences from professionals relying on fabricated citations produced by general AI tools.
Tax figures change every year: thresholds, rates, allowances and limits. A model trained on a mix of years can confidently return last year's number, or blend two years together. The answer looks precise, which is what makes a wrong-year figure so easy to miss.
A subtler hallucination applies a real rule to facts it does not govern. The legislation exists, the relief is real, but it does not apply to the client's situation. This is the hardest kind to catch, because nothing in the answer is fabricated; the error is in the join between rule and facts.
Catching hallucinations is easier once you understand that they are a design consequence, not a glitch.
A general-purpose language model is built to produce plausible, well-formed language. That is what it optimises for. Being correct is a different property, and the two do not always travel together. A model can write a beautifully structured, grammatically perfect paragraph about a relief that simply does not exist, and it will do so with exactly the same confidence as a correct one.
Human experts hedge when they are unsure. General models often do not. An uncertain answer and a certain one can read identically, which strips away the natural signal, a note of doubt, that would otherwise warn a reader to check.
A general model answers from patterns absorbed in training, not from a live document it is reading. Without a real source in front of it, it reconstructs what an answer probably looks like. Most of the time that reconstruction is right; the hallucination is what happens when it is not, and there is no source to catch it. A related design property to watch is what a given tool explicitly will not do, which good tools state openly.
You do not need to be a machine-learning expert to defend against hallucinations. Three disciplined habits catch the large majority.
Never accept a cited case, section or guidance note on trust. Open it and read the relevant part. Confirm it exists, and confirm it says what the answer claims. This single habit catches fabricated authority and most misapplied-rule errors, because reading the source forces you to check the join between rule and facts.
Treat every figure as year-stamped until proven otherwise. For any threshold, rate or allowance, confirm which tax year it belongs to and that it is the right one for the client's period. Wrong-year figures are among the most common and most preventable hallucinations.
Ask whether the cited case or provision actually resolves the question in front of you, or merely touches the topic. An authority can be real, relevant-sounding and still not decide your specific point. Reading it with your exact facts in mind is what separates a genuine answer from a plausible one.
Habits catch hallucinations after the fact. The better strategy is to reduce how many occur in the first place, by choosing tools that are built to answer from sources.
A grounded tool retrieves real source material and answers from it, then shows you that source, rather than generating an answer from memory and hoping it is right. This does not make errors impossible, but it changes their nature: instead of a fabricated citation, you get a real one you can open and check in seconds. It converts the verification tax from an afternoon into a click.
When evaluating any AI tax tool, ask not only how accurate it is but how often it hallucinates, and how that is measured. A published, checkable figure is worth far more than a marketing claim. GAIN Tax publishes a benchmark showing 93.8% correct across 176 questions in 20 UK tax domains, with hallucinated answers held to 1.7%, each answer traceable to its source and kept current. For the full framework on weighing these criteria, see how to choose AI tax research software in the UK.
No tool removes professional responsibility. The point of grounding and low hallucination rates is to make the human's check fast and reliable, not to replace it. The reviewer still opens the source, weighs the judgement and signs underneath.
Hallucinations are the tax profession's version of a well-dressed stranger: confident, plausible and occasionally not who they claim to be. They happen because general models are built for fluency, not truth, and they are most dangerous exactly where tax work is most exacting.
The defence is not to distrust AI, but to use it well. Open every citation, check every tax year, confirm the authority decides your point, and choose tools that answer from real sources and publish how often they get it wrong. Do that, and the answer that sounds right becomes the answer you can prove right.
To see grounded, cited answers on your own questions, create a free GAIN Tax account and try to catch one out.
What is an AI hallucination in tax research?
It is a fluent, confident answer that is wrong: a fabricated or misapplied case, a figure from the wrong tax year, or a real rule applied to facts it does not govern. The danger is that it reads as authoritative.
Why do AI models hallucinate?
Because they are optimised to produce plausible language, not to be correct, and those are different skills. A general model can write a perfect paragraph about a rule that does not exist, with the same confidence as a correct one, especially when answering from memory rather than a live source.
How can I tell if an AI tax answer is hallucinated?
Open every citation and confirm it exists and says what is claimed, check the tax year on every figure, and ask whether the cited authority actually decides your specific point. Most hallucinations fail at least one of those three checks.
Do grounded or cited tools eliminate hallucinations?
No, but they reduce them and change their nature. Instead of a fabricated citation, you get a real source you can open and verify quickly. It makes checking fast and reliable rather than removing the need to check.
What is a good hallucination rate?
Lower is better, and a published, checkable figure matters more than any single number. GAIN Tax reports 1.7% hallucinated answers on a benchmark of 176 UK tax questions across 20 domains, alongside 93.8% correct, with each answer traceable to its source.
Does using a better tool remove my professional responsibility?
No. Grounding and low hallucination rates make your check faster and more reliable, but the professional still opens the source, weighs the judgement and remains accountable for the advice.