"How accurate is it?" is the first question anyone asks about using AI for tax, and it is the right place to start. It is the wrong place to stop. A single accuracy percentage, quoted without context, tells you almost nothing about whether a tool is safe to put in front of a client, because it hides the two things that actually decide that: how the number was measured, and what happens on the answers it gets wrong.
Accuracy for UK tax research is knowable, and it can be high. But the useful question is not "what is the headline number?" It is "what does that number cover, how often does the tool produce a confident wrong answer, and can a human verify each answer from a source?" Those three questions separate a tool you can rely on from one that merely demos well.
This guide explains what accuracy figures really mean, why the hallucination rate matters as much as the headline, how to read a benchmark, and how to judge any AI tax tool on evidence rather than adjectives.
Key takeaways
What an accuracy figure actually means
A percentage on its own is marketing. The same number can be impressive or meaningless depending on what sits behind it.
The method is the message
An accuracy figure is only as good as the test that produced it. Ten easy questions marked generously can yield a flattering percentage that collapses on real work. A large, varied set of genuine tax questions, spread across many domains and marked against primary law, produces a number you can actually trust. Always ask how many questions, across how many domains, and how they were marked.
Breadth versus depth
A tool might be excellent on VAT and weak on capital gains, and a single blended figure hides that. Coverage across domains tells you whether the accuracy is broad or concentrated. GAIN Tax's benchmark spans 21 UK tax domains precisely so the figure reflects range, not a narrow specialism.
Partial credit and honest marking
Real tax answers are not simply right or wrong. Some are substantially correct with a minor gap. An honest benchmark reports partial results separately rather than folding them into either column. GAIN Tax Accuracy Benchmark reports 93.2% correct, 4.8% partially correct and 2.0% incorrect across its 250 questions, so the shape of the performance is visible, not averaged away.
Partially correct means the core answer is right with no critical failure, but the score sits in a middle band because of gaps such as a missing detail or an imprecise citation.
Why the hallucination rate matters more
If you read only one number beyond the headline, read the hallucination rate.
The most expensive failure
A wrong answer that announces itself is cheap; you catch it and move on. A wrong answer that looks right is expensive, because it can pass a busy review and reach a client. The hallucination rate measures how often a tool produces exactly that: a confident, plausible, incorrect answer. Held low, it is the single best indicator that a tool is safe under pressure.
Why grounding lowers it
Tools that answer from retrieved sources, rather than from memory, hallucinate less, because the answer is tethered to a real document. GAIN Tax holds hallucinated answers to 2.0% on its published benchmark, and shows the source behind each answer so a reviewer can confirm it. You can also see how those sources are kept current, since a citation to superseded law is its own kind of error.
Quality beyond right-or-wrong
Accuracy is one dimension; the quality of a correct answer is another. GAIN Tax's benchmark scores answers on legal correctness, calculation, tax-year accuracy and citation quality, each rated close to full marks, so "correct" means correct in the ways that matter to a professional, not merely close.
How to read a benchmark
A benchmark is only useful if you can interrogate it. Here is what to look for.
Look for the denominator
A percentage needs a denominator. 2% hallucination rate means something because it is 2% of 250 stated questions across 21 stated domains. A figure with no question count or domain list is a claim, not a benchmark.
Look for the failure breakdown
A trustworthy benchmark shows its wrong answers, not just its right ones. The split between partially correct and outright incorrect, and the hallucination rate specifically, tells you how the tool fails, which matters more than how often it succeeds.
Look for independence and repeatability
Note who marked the answers and against what. And remember that a different set of questions can produce a different result, which is why a good benchmark is transparent about its method and honest about its limits. GAIN Tax states this openly alongside its benchmark and its responsible-use limits.
Judging a tool on evidence, not claims
With the concepts in place, evaluating a tool becomes practical.
Test it on your own questions
The most honest benchmark is your own. Put real questions from your practice to any tool, including the awkward ones, and judge the answers and their sources yourself. A tool that grounds each answer in law you can open makes this quick.
Weigh accuracy against verifiability
A slightly lower headline accuracy with full source citation can be safer in practice than a higher figure you cannot check, because the checkable tool lets you catch the errors it does make. Verifiability is what turns a good accuracy figure into a defensible workflow.
Consider proof beyond the number
Accuracy figures sit alongside other evidence: what firms actually say in reviews, how the tool handles your workflow, and whether the vendor is transparent about limits. For the full decision framework, our pillar guide walks through how to choose AI tax research software in the UK end to end.
Conclusion
AI can be highly accurate for UK tax research, and the best tools now answer the large majority of real questions correctly. But accuracy is a starting point, not a verdict. What makes a tool safe to rely on is the method behind its number, a low rate of confident wrong answers, and the ability for a human to verify each answer from a source in seconds.
Judge tools on evidence you can check: a published benchmark with a real denominator, a stated hallucination rate, and cited answers you can open. GAIN Tax reports 93.2% correct across 250 questions in 21 UK tax domains, with hallucinated answers held to 2.0% and every answer traceable to its source, precisely so you can judge it rather than take it on faith.
Try it on the questions that matter to you. Create a free GAIN Tax account, or book a call to see it against your firm's own research.
Frequently asked questions
How accurate is AI for UK tax research?
The best tools now answer the large majority of real questions correctly. GAIN Tax reports 93.2% correct across 250 questions in 21 UK tax domains on its published benchmark. But an accuracy figure is only meaningful with its method and its hallucination rate alongside it.
Why does the hallucination rate matter more than accuracy?
Because a confident wrong answer is the costly failure. It survives a quick read and can reach a client. The hallucination rate measures how often that happens; held low, it is the best single indicator that a tool is safe under pressure. GAIN Tax reports 2.0%.
What makes a benchmark trustworthy?
A real denominator (how many questions, across how many domains), a breakdown of how it fails as well as how it succeeds, transparency about method and marking, and honesty that a different question set can give a different result.
Can I trust an AI tax answer without checking it?
No. The point of high accuracy and low hallucination is to make your check fast and reliable, not to remove it. A tool that shows its source lets you confirm each answer in seconds, but the professional remains responsible for the advice.
How should I evaluate an AI tax tool myself?
Test it on your own real questions, including the awkward ones, and judge both the answers and their sources. Weigh headline accuracy against whether you can verify each answer, and look at wider proof such as firm reviews and vendor transparency about limits.
Does higher accuracy always mean a better tool?
Not necessarily. A slightly lower headline accuracy with full, checkable citations can be safer than a higher figure you cannot verify, because you can catch the errors it does make. Verifiability turns an accuracy figure into a defensible workflow.

