AI Tax Tool Testing: How to Build a Test File for Your Own Work
Every few weeks a new evaluation appears showing that one model is some percentage better than another. The number is usually produced by a vendor, on questions the vendor chose, with errors the vendor planted, and without independent replication. Read as marketing it is unremarkable. Read as evidence for a purchasing decision it is close to useless, because it measures a task that resembles nobody's actual job.
A practice does not need a better published benchmark. It needs twenty of its own questions, the answers it already knows, and an afternoon.
What follows is the method, not a score. It works against any tool, costs nothing but time, and produces something no vendor can hand you: evidence about the work your firm actually does.
Key takeaways
Why published evaluations answer the wrong question
They measure the author's task
An evaluation is a set of questions plus a scoring rule. Both belong to whoever built it. When a vendor reports a large improvement on one task and a modest one on average, the honest reading is that a particular task moved, as measured by a party with an interest in the result and without outside replication.
None of that makes the number dishonest. It makes it inapplicable. Your firm's question mix is not the author's question mix.
The distribution problem
Most published evaluations sample broadly across a subject. A practice does not work broadly. It works on a narrow band of recurring situations, determined by its client base: a market town firm full of farms and furnished holiday lets, a City practice full of share schemes, a bookkeeping practice where nine questions in ten are VAT liability on borderline supplies.
A tool can perform respectably across a broad sample and fail precisely where your fee income sits.
⚠️ Watch: Treat any evaluation you did not design as evidence that a tool works on somebody's questions. It carries no information about yours until you have tested it on yours.
What a good published benchmark is actually for
A well-constructed public benchmark, with its method disclosed, has one honest use: eliminating candidates. A tool that performs poorly on a disclosed, independent test is unlikely to surprise you favourably on your own. It earns a place on the shortlist, never a place in the workflow. Our own benchmark methodology is published for that purpose, including the failures, and it scores every answer on the same two questions this article asks: is it correct, and does the source it cites support it. We would still rather a firm ran its own file.
Building the file: twenty questions from closed work
Where the questions come from
Closed files, not imagination. Take the last two years of technical queries the practice actually resolved, whether by research, by a call to a helpline, or by an argument between two partners.
Closed work has the property that matters: you already know the answer, and you know it was checked.
The mix
A workable twenty breaks down roughly as follows.
Writing a question properly
Write it as you would put it to a colleague, with the facts a colleague would need and nothing more. Resist tidying the facts into a textbook shape. The messiness is the point, because messy is what arrives.
Record the known-correct answer beside it, with the authority you relied on at the time. That pairing is the whole test.
✅ In practice: Write the answers before you run anything. An answer written after seeing a tool's output is no longer an independent standard, and the temptation to soften it is stronger than anyone expects.
The four questions that separate tools
1. Does the citation resolve?
Take every citation in an answer and open it. Does the section say what the answer claims it says? Does the paragraph number exist? Does the case concern the point it is cited for?
This single check disposes of more candidates than any other, and it requires no judgement about tax at all.
2. Does it know which year it is in?
A rate, a threshold or an allowance is only correct relative to a tax year. An answer that gives a figure without anchoring it to a year has failed, even when the figure happens to match the current one.
3. What does it do at the edge of its competence?
Feed it the two questions with no clean answer. The useful response describes the competing positions and names the uncertainty. The dangerous response picks one and writes it up beautifully.
"How a tool behaves when it does not know is more informative than how it behaves when it does."
4. Does it repeat a superseded rule under pressure?
This is what the planted traps are for. Ask about a relief that changed, or a threshold that moved, and see whether the old answer arrives dressed in the same confidence as the new one.
Scoring without pretending to be a laboratory
Keep it coarse
Three outcomes are enough: usable as drafted, usable after correction, or wrong. Finer scales invite argument and add nothing to the decision.
A fourth category is worth recording separately: wrong with a citation. An error carrying an authority that appears to support it is more dangerous than a bare error, because it survives a casual review. Stanford's 2025 study of AI legal research tools, the closest peer-reviewed evaluation of this class of tool, calls such an answer misgrounded and counts it as a hallucination in its own right, alongside plain false statements. Its authors' point is that a working link is the weakest possible test of a citation.
Who marks it
Someone who did not write the questions, wherever the firm is large enough to allow it. The person who wrote a question knows what they meant, and unconsciously grades against their intention rather than against what they actually asked.
What a pass looks like
There is no universal threshold, and any vendor offering one is guessing at your risk appetite. The practical test is narrower: on the eight bread-and-butter questions, would you have sent that answer to a client after a normal review? If the answer is no more than six times out of eight, the tool is not yet saving anyone time.
Running it again, and what to do with the results
A test file written once and never re-run measures a product on a single day. Models change, corpora are updated, and guidance moves underneath both.
- Re-run the same twenty quarterly. Same questions, same answers, same marker where possible.
- Add the year's new traps. Every Budget and every Finance Act creates fresh superseded answers. Two new questions a year keeps the file honest.
- Log the misses, not the score. A percentage is for reporting. The list of what it got wrong is what tells you which questions still need a human.
- Keep the file when you change tools. It is the only asset in this exercise, and it makes the next evaluation an afternoon instead of a project.
- Record it as part of your AI governance. A dated test file, with results, is evidence of the reasonable care a firm took before relying on a tool, which is a question professional bodies are increasingly asking.
Our limitations and responsible use page sets out where we think the boundary of tool reliance sits, and the sources and update policy describes how the underlying material is kept current. If you want a subject for the first run, create a free account and put your own twenty questions to it before you put them to anybody else.
Conclusion
The published-benchmark conversation has trained buyers to compare numbers that were produced elsewhere, on tasks chosen by someone else, for reasons that have nothing to do with a particular practice. Meanwhile the evidence a firm actually needs is sitting in its own closed files, already checked, costing nothing to assemble.
Twenty questions. Answers written first. Citations opened one by one. Three planted traps and two questions with no clean answer. An afternoon, once, and an hour a quarter after that.
It produces a smaller, duller, far more useful number than anything in a press release, and it is the only number that describes your risk.
For the wider criteria worth applying before a tool reaches the test file at all, see our guide to choosing AI tax research software in the UK. Firms evaluating across several seats can book a call, and common objections are collected in the FAQ.
Frequently asked questions
How many questions do I need to test an AI tax tool?
Twenty drawn from closed files is enough to separate a usable tool from an unusable one. The value comes from the mix rather than the volume: everyday questions from your own client base, a few genuinely hard ones, some date-sensitive figures, and deliberate traps built on rules that have since changed.
Why not rely on a published AI benchmark?
A published benchmark measures the questions its author selected, scored by a rule its author wrote. Where the author is also the vendor and no independent replication exists, it describes a chosen task rather than yours. Use disclosed, independent benchmarks to eliminate candidates, then test the survivors on your own work.
What is the single most useful check on an AI tax answer?
Opening every citation. Confirm the provision exists, says what the answer claims, and is current for the year in question. This requires no tax judgement, catches the most dangerous class of error, and disposes of weak tools faster than any other test.
What is a "planted trap" question?
A question whose correct answer changed at a known point, such as a threshold that moved or a relief that was withdrawn. It tests whether a tool supplies a superseded answer with the same confidence as a current one, which is the failure mode most likely to reach a client unnoticed.
Who should mark the test in a practice?
Someone other than the person who wrote the questions, where the firm is large enough. Question authors grade against what they meant rather than what they asked. Marking should also be coarse: usable as drafted, usable after correction, or wrong, with errors that carry a plausible citation recorded separately.
How often should a test file be re-run?
Quarterly, using the same questions and the same known answers, with two new trap questions added each year as rules change. Re-running matters because models, corpora and guidance all move; a single evaluation describes a product on one day rather than over a period of reliance.

