Every AI-assisted tax answer arrives in the same confident voice, whether it is right or wrong. That is the whole problem in one sentence. A model has no tell: the paragraph that quotes last year's threshold reads exactly like the paragraph that quotes this year's.
For a profession where the signature at the bottom carries the liability, that makes verification a workflow question rather than a philosophical one. Whatever drafts the answer, the professional who signs it owns it – so the only durable defence is a routine dull enough to be followed under deadline pressure.
This guide sets out that routine: why fluency and accuracy are unrelated, a five-step verification pass, the specific traps that hide one layer below a plausible answer, and what "designed to be checked" looks like in a tool.
Key takeaways
-
Tone tells you nothing. A generative model produces smooth, authoritative text regardless of whether the underlying answer is right.
-
Verify against the primary source, not against the model's summary of it. Retrieve, check currency, match the facts, test the exception, record the trail.
-
The dangerous errors are the near-misses: a real relief with a missing condition, last year's threshold, a real case that decides something else.
-
Publish-grade tools show their sources. If a claim cannot be traced in seconds, verification gets skipped in practice, whatever the policy says.
-
Accountability does not transfer. The professional who signs the work remains responsible for it, which is why checkability is a procurement criterion, not a nice-to-have.
Why fluency is not accuracy
A language model is optimised to produce plausible text. Plausibility and correctness overlap most of the time, which is exactly what makes the gaps dangerous: a tool that is right nine times in ten trains you to stop looking on the tenth.
The confidence problem
Human experts signal uncertainty. They hedge, they say "I'd want to check that", they pause. A model does none of this unless it has been specifically built to. The absence of hedging is not evidence of reliability; it is a property of the format.
What the numbers actually look like
Measured accuracy on real professional questions is high but not total, and honest vendors publish the figure. On our own benchmark, GAIN Tax answered 93.8% of 176 real UK tax questions correctly, with 4.5% partially correct and 1.7% incorrect or hallucinated, across 20 domains. Every question and answer was validated by a qualified tax specialist under a pre-registered protocol, and the full benchmark is published including the failures.
A 1.7% error rate is good. It is also not zero, and on a book of several hundred queries a year, "not zero" is the entire argument for having a checking routine.
The five-step verification pass
1. Retrieve the actual source
Open the source itself – gov.uk, legislation.gov.uk, the HMRC manual page, the judgment – rather than relying on the model's description of it. A summary can be faithful and still omit the sentence that decides your client's case.
2. Confirm it is current
Check the effective date, and check whether a later measure has superseded it. Tax content ages badly and asymmetrically: rates and thresholds move on predictable dates, but consultations, cases and enforcement powers move whenever they move.
3. Match the facts, line by line
Hold each figure, threshold, date and condition against the source. Not the gist – the specific numbers. This is where near-miss errors surface, and it takes a couple of minutes.
4. Test the exception
Ask whether your client's facts sit inside a special rule rather than the general one. Most professional value lives here, and most general answers assume the general case without saying so.
5. Record the trail
Note the source and the date on the file, so the reasoning survives the person who did it. A position that cannot be reconstructed in eighteen months is not defensible, however sound it was at the time.
Quick checklist
- Have I opened the primary source, not a summary?
- Is it current at the date of the advice?
- Does every figure in the answer match the source?
- Have I asked whether an exception applies to these facts?
- Is the source and date recorded on the file?
The traps that hide one layer down
The plausible citation
A reference that looks entirely normal – correct format, plausible name, plausible year – to a case or manual paragraph that does not exist, or does not say what is claimed. It survives a glance and fails a click.
The almost-right threshold
Last year's number wearing this year's date. This is the most common and least dramatic failure, and it is why step 2 is separate from step 1.
The wrong-holding case
A real, correctly-cited case that decides a different question. The citation checks out, the name is right, and the proposition it is offered for is not what the court held.
The confidently general answer
Technically correct as a statement of the general rule, and wrong for a client sitting inside an exception. No amount of source-checking catches this one; only asking the exception question does.
What "designed to be checked" looks like
Verification that depends on willpower does not survive a busy week. It has to be built into the tool, which gives a practice a short list of procurement questions:
- Is every claim one click from its primary source? Not "trained on tax data" – an actual, openable reference for the specific sentence.
- Is it grounded in UK legislation and HMRC material, rather than the open web?
- Does it leave an audit trail a colleague or a regulator could follow?
- Can you see when the underlying material was last updated?
- Does it behave well when unsure, flagging doubt rather than producing a tidy, confident answer?
This is the thinking behind how GAIN Tax is built: UK-grounded answers with the source attached to every claim, so checking takes seconds instead of being skipped. Our sources and update policy sets out what the system reads and how currency is maintained, and our limitations and responsible use page is deliberately explicit about what the tool will not do. If you are evaluating options for a practice, our guide to choosing AI tax research software covers the wider criteria.
Building the routine into a practice
- Make verification a step in the workflow, not a virtue. If the file template has a "source and date" field, it gets filled in.
- Decide where AI is allowed and on what work. Research and first drafts are low-risk, high-value ground; filing positions and final advice are not.
- Train juniors on the traps, not just the tool. The plausible citation and the wrong-holding case are learnable patterns.
- Keep the exception question explicit in review notes, because it is the one no automated check will do for you.
Conclusion
The value of an AI research tool is not that it removes the checking. It is that it moves the expensive part of research – finding the right source – from twenty minutes to twenty seconds, leaving the professional judgement where it belongs.
Fluent text is the perfect hiding place for a near-miss. A short, boring routine, applied every time, is what stands between a smooth paragraph and a client file you would be content to defend.
Want a tool where every claim is one click from its source? Start a free trial of GAIN Tax, or book a call to talk through a firm-wide rollout. More practical guides are on the GAIN Tax blog.
Frequently asked questions
Can I rely on an AI answer for tax advice? Not without checking it. The professional who signs the work remains responsible for its accuracy, so an AI answer is a draft until the primary source behind each claim has been opened and matched. Used that way it saves substantial time.
How do I spot an AI hallucination in tax research? Open the citation. Hallucinated references usually look entirely normal and fail on the click: the case does not exist, the manual paragraph says something else, or the figure is last year's. Checking currency separately from existence catches most of the rest.
How accurate is AI on UK tax questions? It depends on the tool and how it is grounded. On our published benchmark, GAIN Tax answered 93.8% of 176 real UK tax questions correctly, 4.5% partially, and 1.7% incorrectly, across 20 domains, with every answer reviewed by a qualified specialist.
What are the most common errors in AI tax answers? Near-misses rather than nonsense: a real relief with a condition omitted, a threshold from the previous tax year, a genuine case cited for a proposition it does not support, and correct general rules applied to facts that sit inside an exception.
Does using AI change my professional responsibility? No. Responsibility for the accuracy of the work stays with the professional who signs it, whatever produced the draft. That is why the ability to check a tool's output quickly is a procurement criterion rather than a preference.
How long should verifying an answer take? With a tool that links each claim to its primary source, a straightforward answer takes a couple of minutes to verify. The time cost sits in retrieving sources, which is exactly the part good grounding removes.

