Most AI software reaches accountancy practices wearing a cost story. Fewer hours, lower cost per file, a saving you can put in a board paper.
Then you look at what the market actually reports. In an Accounting Talent Index survey of around 500 UK firms published in June 2026, 73% said they had turned away work because they did not have the staff. In the same survey, 74% thought sustained workloads could push people out of the profession altogether. The survey was commissioned by an outsourcing provider, which is worth stating plainly, and it still describes something most principals recognise immediately.
A firm turning away work is not primarily trying to make its existing work cheaper. It is trying to say yes more often. Those two problems are bought in completely different ways, and modelling the second one as though it were the first produces a business case nobody believes.
This article sets out how to model the return properly, what to measure, and the three questions that decide whether a tool releases capacity or merely relocates work.
Key takeaways
Why the cost model understates the case
It measures the wrong scarce resource
In a firm at capacity, the constraint is not the price of an hour. It is the existence of an hour belonging to somebody senior enough to sign the work. Saving money on a resource you are not short of produces a small, honest, unexciting number.
It ignores the revenue that never appeared
Work declined in March does not show up anywhere in the accounts. There is no line for it, no variance to explain, and no meeting about it. It is the largest number in the analysis and the only one that is entirely invisible.
It flatters the wrong tools
A pure cost model rewards whichever tool is cheapest per query. A capacity model rewards the tool that produces an answer a reviewer can sign off quickly, which is a different property and often a more expensive one.
The metric that actually decides it
Time to a defensible answer
Measure from the moment the question is asked to the moment somebody is willing to put it in front of a client under the firm's name. Include:
- the time to produce the answer;
- the time to open and read the cited sources;
- the time to confirm the tax year or period;
- the time to confirm the authority decides the point rather than merely mentioning it;
- any time spent reconstructing where an answer came from.
That last line is where projected savings go to die. An answer produced in ten seconds and verified over forty minutes has cost you forty minutes and ten seconds.
"Checking a citation takes a minute. Reconstructing where an answer came from takes an afternoon."
Run the measurement on your own questions
Vendor benchmarks tell you about a vendor's question set. Your own recurring questions tell you about your practice. Pick ten that come up genuinely often, put them through any tool you are assessing, and time the whole loop including review.
✅ In practice: Ten of your own questions, timed end to end, will tell you more in an afternoon than any demonstration. You can create a GAIN Tax account and run exactly that test, and our published benchmark sets out the method behind our own figures, including how failures are counted.
Three questions that decide whether capacity is released
1. Is the source shown?
An answer with the legislation, guidance or case pinned to the point turns verification into a click. An answer without one turns it into research. This single property does more to the end-to-end time than raw accuracy does.
2. What is the failure rate, and will they show you?
Every vendor publishes an accuracy figure. Fewer publish the shape of their failures. Ask three things: who wrote the questions, what counts as a failure, and can you see the ones it got wrong.
Our own run: 250 questions across 21 UK tax domains, 233 correct, 12 partially correct, 5 incorrect. A tax specialist writes each question and the golden answer it is marked against, validates every result and signs off the report, and each response is scored on eight dimensions against a fixed rubric with three marked critical.
The partially correct answers deserve the most attention. An answer that lands roughly right and fumbles a step is the one most likely to pass a tired reviewer.
3. Does it fit the work you would actually take on?
Capacity released in an area you do not sell is not capacity. Check coverage against the work you would accept if you had the hours: the recurring compliance questions, the arguable positions, the areas where you currently refer out.
What not to model
Three inputs turn a business case into fiction, and all three are common.
- A blanket percentage applied to every hour in the practice. Research and first drafting are a slice of the working week, not the whole of it. Applying a research-time saving to total chargeable hours inflates the answer by an order of magnitude.
- Savings on fixed-fee work you were already delivering. Compressing the effort on a job billed at a fixed fee improves the margin on that job, which is real but modest, and it is a different number from the fee on work you could newly accept.
- Headcount reductions. A firm that cannot recruit is not about to release staff. Modelling a licence against a salary in a market where 73% of firms are turning work away describes a practice nobody in this survey is running.
Strip those out and what remains is usually small, credible, and considerably easier to defend to a partner group than a spreadsheet promising a transformation.
Modelling the return without inventing numbers
Start with one engagement, not a percentage
The credible version of this business case is narrow. Identify one type of work the practice declined in the last year for want of hours. Estimate the fee. Ask what proportion of the effort was research and first drafting, and how much of that a tool with shown sources would compress.
If recovering one such engagement covers the annual licence cost, the case is made and everything beyond it is upside. If it does not, no spreadsheet of percentage savings applied across every hour in the firm will rescue it.
A worked shape
⚠️ Watch: Do not model a saving on hours you would not have billed anyway. The return on a capacity tool comes from work you take on, not from theoretical minutes shaved off work you were already doing at a fixed fee.
For firm-level decisions
Multi-seat purchases change the arithmetic, because the constraint moves from individual research time to review throughput across the practice. That is a conversation worth having properly rather than modelling in a spreadsheet: book a call if it is the position you are in, and what other firms say is a reasonable place to start.
Conclusion
The capacity constraint reported across UK practices is real, and it is not the same problem as cost. A firm declining work because there is nobody to give the file to gains nothing meaningful from a cheaper hour, and gains a great deal from being able to say yes.
That makes the buying question narrow and answerable. Of the work you turned away last year, which slice would you have accepted if the research and the first draft arrived already sourced and ready to be reviewed rather than rebuilt?
Answer that with one engagement, one fee and one honest timing test, and the decision usually makes itself. The wider framework for assessing tools sits in our guide to choosing AI tax research software in the UK, and there is more on the blog. To run the test on your own questions, create an account.
Frequently asked questions
How do you calculate ROI on AI tax research software?
Model it on recovered work rather than on hours saved. Identify one engagement the practice declined for lack of capacity, estimate the fee, work out what share of the effort was research and first drafting, and time how much of that a tool compresses on your own questions. If one recovered engagement covers the annual licence, the case is made.
Is AI tax research worth it for a small practice?
It depends on whether the practice is constrained by cost or by capacity. A firm with spare capacity and fee pressure is buying a cost saving, which is usually modest. A firm turning away work for lack of hours is buying the ability to accept it, which is where the material return sits.
What should I measure when trialling a tax research tool?
Time to a defensible answer, measured end to end and including the review: producing the answer, opening and reading the cited sources, confirming the tax year, and confirming the authority decides the point. An answer produced instantly and verified over forty minutes has saved nothing.
How accurate is GAIN Tax?
On our published benchmark, 93.2% correct across 250 questions in 21 UK tax domains, with 4.8% partially correct and 2.0% incorrect. A tax specialist writes every question and golden answer, validates each result and signs off the report, and responses are scored on eight dimensions with three marked critical. The full method is on our benchmark page.
Why does showing the source matter more than accuracy?
Because verification time dominates the end-to-end figure. Above a certain accuracy threshold, another decimal point changes very little, while an answer that shows its source turns a reviewer's check from research into a click. That difference is what decides whether capacity is genuinely released.
How many questions should I test before deciding?
Ten of your own recurring questions, timed end to end, is enough to separate tools in an afternoon. They should be questions your practice actually meets rather than difficult edge cases, because the return comes from routine work compressing reliably rather than from exotic questions being answered impressively.

