Every AI-assisted tax answer arrives in the same confident voice, whether it is right or wrong. That is the whole problem in one sentence. A model has no tell: the paragraph that quotes last year's threshold reads exactly like the paragraph that quotes this year's.
For a profession where the signature at the bottom carries the liability, that makes verification a workflow question rather than a philosophical one. Whatever drafts the answer, the professional who signs it owns it – so the only durable defence is a routine dull enough to be followed under deadline pressure.
This guide sets out that routine: why fluency and accuracy are unrelated, a five-step verification pass, the specific traps that hide one layer below a plausible answer, and what "designed to be checked" looks like in a tool.
Tone tells you nothing. A generative model produces smooth, authoritative text regardless of whether the underlying answer is right.
Verify against the primary source, not against the model's summary of it. Retrieve, check currency, match the facts, test the exception, record the trail.
The dangerous errors are the near-misses: a real relief with a missing condition, last year's threshold, a real case that decides something else.
Publish-grade tools show their sources. If a claim cannot be traced in seconds, verification gets skipped in practice, whatever the policy says.
Accountability does not transfer. The professional who signs the work remains responsible for it, which is why checkability is a procurement criterion, not a nice-to-have.
A language model is optimised to produce plausible text. Plausibility and correctness overlap most of the time, which is exactly what makes the gaps dangerous: a tool that is right nine times in ten trains you to stop looking on the tenth.
Human experts signal uncertainty. They hedge, they say "I'd want to check that", they pause. A model does none of this unless it has been specifically built to. The absence of hedging is not evidence of reliability; it is a property of the format.
Measured accuracy on real professional questions is high but not total, and honest vendors publish the figure. On our own benchmark, GAIN Tax answered 93.8% of 176 real UK tax questions correctly, with 4.5% partially correct and 1.7% incorrect or hallucinated, across 20 domains. Every question and answer was validated by a qualified tax specialist under a pre-registered protocol, and the full benchmark is published including the failures.
A 1.7% error rate is good. It is also not zero, and on a book of several hundred queries a year, "not zero" is the entire argument for having a checking routine.
Open the source itself – gov.uk, legislation.gov.uk, the HMRC manual page, the judgment – rather than relying on the model's description of it. A summary can be faithful and still omit the sentence that decides your client's case.
Check the effective date, and check whether a later measure has superseded it. Tax content ages badly and asymmetrically: rates and thresholds move on predictable dates, but consultations, cases and enforcement powers move whenever they move.
Hold each figure, threshold, date and condition against the source. Not the gist – the specific numbers. This is where near-miss errors surface, and it takes a couple of minutes.
Ask whether your client's facts sit inside a special rule rather than the general one. Most professional value lives here, and most general answers assume the general case without saying so.
Note the source and the date on the file, so the reasoning survives the person who did it. A position that cannot be reconstructed in eighteen months is not defensible, however sound it was at the time.
A reference that looks entirely normal – correct format, plausible name, plausible year – to a case or manual paragraph that does not exist, or does not say what is claimed. It survives a glance and fails a click.
Last year's number wearing this year's date. This is the most common and least dramatic failure, and it is why step 2 is separate from step 1.
A real, correctly-cited case that decides a different question. The citation checks out, the name is right, and the proposition it is offered for is not what the court held.
Technically correct as a statement of the general rule, and wrong for a client sitting inside an exception. No amount of source-checking catches this one; only asking the exception question does.
Verification that depends on willpower does not survive a busy week. It has to be built into the tool, which gives a practice a short list of procurement questions:
This is the thinking behind how GAIN Tax is built: UK-grounded answers with the source attached to every claim, so checking takes seconds instead of being skipped. Our sources and update policy sets out what the system reads and how currency is maintained, and our limitations and responsible use page is deliberately explicit about what the tool will not do. If you are evaluating options for a practice, our guide to choosing AI tax research software covers the wider criteria.
The value of an AI research tool is not that it removes the checking. It is that it moves the expensive part of research – finding the right source – from twenty minutes to twenty seconds, leaving the professional judgement where it belongs.
Fluent text is the perfect hiding place for a near-miss. A short, boring routine, applied every time, is what stands between a smooth paragraph and a client file you would be content to defend.
Want a tool where every claim is one click from its source? Start a free trial of GAIN Tax, or book a call to talk through a firm-wide rollout. More practical guides are on the GAIN Tax blog.
Can I rely on an AI answer for tax advice? Not without checking it. The professional who signs the work remains responsible for its accuracy, so an AI answer is a draft until the primary source behind each claim has been opened and matched. Used that way it saves substantial time.
How do I spot an AI hallucination in tax research? Open the citation. Hallucinated references usually look entirely normal and fail on the click: the case does not exist, the manual paragraph says something else, or the figure is last year's. Checking currency separately from existence catches most of the rest.
How accurate is AI on UK tax questions? It depends on the tool and how it is grounded. On our published benchmark, GAIN Tax answered 93.8% of 176 real UK tax questions correctly, 4.5% partially, and 1.7% incorrectly, across 20 domains, with every answer reviewed by a qualified specialist.
What are the most common errors in AI tax answers? Near-misses rather than nonsense: a real relief with a condition omitted, a threshold from the previous tax year, a genuine case cited for a proposition it does not support, and correct general rules applied to facts that sit inside an exception.
Does using AI change my professional responsibility? No. Responsibility for the accuracy of the work stays with the professional who signs it, whatever produced the draft. That is why the ability to check a tool's output quickly is a procurement criterion rather than a preference.
How long should verifying an answer take? With a tool that links each claim to its primary source, a straightforward answer takes a couple of minutes to verify. The time cost sits in retrieving sources, which is exactly the part good grounding removes.