For three years the story of AI in finance has been a race for accuracy. Each new model scored higher on the benchmarks, wrote more fluent prose and answered more questions correctly, and the assumption was that if accuracy kept climbing, trust would follow. In 2026 the data told a more interesting story.
Research published this year found that most finance leaders would turn down a highly accurate AI if it could not show its work. The number they will not accept without is not another decimal point of accuracy. It is an explanation. That single finding reframes what "good AI" means in a professional setting, and it has a direct bearing on how tax research tools should be chosen.
This article explains what the verification tax is, why explainability now matters more than the last few points of accuracy, and what an adviser should demand from any AI tool that touches a client file.
The phrase comes from an IDC white paper, sponsored by Sage and based on a survey of more than two thousand senior finance decision-makers. It names a cost that most teams feel but few had measured.
When an AI produces an answer that a professional must stand behind, someone has to check it. That checking is real work: opening sources, reconciling figures, confirming that a cited authority says what the model claims. The research found finance teams spend, on average, around thirteen hours a week on this reconstruction and validation, with nearly half spending fifteen hours or more, and a smaller group spending thirty hours and up.
The promise of AI is time saved. The verification tax is the portion of that saving that leaks straight back out. The same research estimated that a meaningful share of expected productivity gains is lost to reverse-engineering outputs so they can be explained to stakeholders. A tool that answers quickly but opaquely does not remove the work; it relocates it, from producing the answer to proving it.
Giving the cost a name matters because it changes the buying question. A tool evaluated only on speed and accuracy looks impressive in a demo. The same tool evaluated on total time to a defensible answer, including the checking, can look very different.
It sounds counter-intuitive that professionals would decline a more accurate tool. It makes complete sense once you separate two things that are usually blurred together.
Accuracy measures how often the model is right. Explainability measures how quickly a human can confirm that it is right. In professional work the second property carries the risk, because a partner signs their name under the conclusion, not under the model's confidence score.
Above a certain threshold, extra accuracy buys little. The difference between a model that is right 97% of the time and one that is right 98% of the time is real but small, and a human has to check the output either way, because they cannot know in advance which answer falls in the wrong percent. What actually reduces the checking burden is not a higher score. It is a visible source.
Tax advice is regulated, and a professional who relies on an answer they cannot explain is exposed if it is wrong. A chain is only as strong as its weakest link, and in an AI workflow the weakest link is the moment a person has to trust an output they cannot trace. Explainability strengthens exactly that link.
The practical difference between an opaque and a transparent tool is easiest to see in the checking step.
With an unexplained answer, verification means rebuilding the reasoning from scratch: finding the legislation, the case, the correct tax year, and confirming each one. With a cited answer, verification means opening the source the tool already provided and reading it. The first can take the better part of an hour on a knotty point. The second takes a couple of minutes.
When two colleagues disagree about an answer, a citation turns the argument into a shared task: open the source and read it together. Without a source, the disagreement stays a contest of confidence. This is why professional tools that ground each answer in primary material are easier to adopt across a team.
A source is only useful if it is current. Tax law changes constantly, and an answer citing superseded guidance is worse than no answer, because it looks authoritative. A serious tool is explicit about where its answers come from and how the law is kept current, and honest about what it will not do.
If the real cost of AI is the checking, then the buying criteria should target the checking directly.
In any evaluation, put a genuine question to the tool and look at what comes back with the answer. Does it provide the specific legislation or case, pinned to the point? Can your reviewer open that source and confirm the answer in minutes? A tool that hands over a source is buying back the verification tax; one that hands over only prose is charging it.
Judge tools on the whole loop, not just the first response. Time how long it takes to reach an answer you would put in front of a client, including the check. The fastest first draft is not always the fastest defensible answer.
A hallucinated answer wearing a citation is the most expensive failure, because it survives a quick read. Ask how the tool measures and reports wrong answers. GAIN Tax publishes a benchmark showing 93.2% correct across 250 questions in 21 UK tax domains, with hallucinated answers held to 2.0%, and every answer traceable to its source. Published, checkable numbers beat marketing adjectives every time.
The best tool for casual questions is not necessarily the best for client-facing research. For a full framework on matching an AI research tool to how your practice actually works, see our guide on how to choose AI tax research software in the UK.
The market has quietly moved the goalposts. For years the question was how accurate an AI could be. The 2026 research shows that professionals have already answered a different question: how much of my time does this tool actually save once I have made its answers safe to use?
The tools that win the next phase are the ones that treat verification as the product, not an afterthought. They show their sources, keep them current, and let a human confirm each answer in a click rather than an afternoon. An answer you can trace is worth ten you can only admire.
To see what a cited, benchmarked answer looks like on your own questions, create a free GAIN Tax account and put a real one to the test.
What is the verification tax?
It is the human effort spent reconstructing, checking and defending AI output before a professional can rely on it. Research by IDC, sponsored by Sage in 2026, put it at around thirteen hours a week on average for finance teams, with nearly half spending fifteen hours or more.
Why would anyone reject a more accurate AI?
Because accuracy and explainability are different things. A more accurate model that cannot show its reasoning still has to be checked, and the checking is where the cost sits. The IDC research found 71% of finance leaders would reject a 99%-accurate tool that could not explain itself.
Does explainability replace the need to check answers?
No. It makes checking far faster. A cited answer turns verification from rebuilding the reasoning yourself into opening the source the tool provided and reading it. The professional still confirms the answer and remains responsible for it.
What is a hallucination in this context?
A hallucination is a fluent, plausible answer that is wrong: a misapplied rule, a case cited for something it does not say, or a figure from the wrong year. The danger is that it reads as authoritative, which is why the hallucination rate matters as much as headline accuracy.
How can I measure the verification tax in my own practice?
Time the full loop on real questions: from asking the tool to reaching an answer you would put in front of a client, including the checking step. Compare tools on total time to a defensible answer, not just the speed of the first response.
Is a published benchmark better than a vendor accuracy claim?
A published, checkable benchmark that shows the number of questions, the domains covered and the hallucination rate is far more useful than an unsupported percentage. It lets you judge the method, not just the headline. GAIN Tax's benchmark reports 93.2% correct across 250 questions in 21 UK tax domains, with 2.0% hallucinated.