GAIN Tax Accuracy Benchmark

UK Tax AI Performance Benchmark

A technical accuracy benchmark for AI tax research: how well the GAIN Tax model performs on real-world UK tax questions, covering tax knowledge and judgment, calculations, and applied scenarios. It is built from real questions and situations faced by accountants, tax advisers and other tax professionals.

Results
176 questions
20 UK tax domains
93.8 %165 Correct
4.5 %8 Partially correct
1.7 %3 Incorrect (Hallucinated)
Result Classifications

How each answer is classified

Each scored answer is correct, partially correct, or incorrect, and incorrect answers stay in the published spread. Every result is validated by our tax specialist under a pre-registered protocol.

Correct

An answer is correct when it clears the weighted threshold with no failure on the three critical dimensions: law, tax year and calculation.

Partially correct

The core answer is right with no critical failure, but the score sits in a middle band because of gaps such as a missing detail or an imprecise citation.

Incorrect (hallucinated)

The answer is wrong in law, miscalculated, on the wrong tax year, misgrounded, or fabricated. A failure on any critical dimension, or a generally wrong answer.

Reliability

Hallucinations, defined plainly

A hallucination is when the model says something that is not true, or claims a source backs up a point when it does not. We count an answer as a hallucination on any of three grounds: fabricated content (an invented rule, rate, section or citation), misgrounded citation (a real-looking source that does not support the claim), or contradicted content (a figure or position the Golden answer states differently). A correct answer that simply lacks a citation is not a hallucination.

When the system holds no relevant sources for the topic in its database, it declines and says so rather than inventing an answer. This is the safe behaviour and the opposite of a hallucination. None occurred on this question set.

Hallucination rate1.7%

Fabricated, misgrounded, or contradicted. The full field-aligned figure, computed from the validated results. The same 3 answers recorded as incorrect.

Quality dimensions

The dimensions behind the score

Every answer is graded on eight dimensions, validated by our tax specialist. These four carry the most weight. The full rubric sits in the methodology below.

Legal correctness4.85 /5

Whether the substance of the answer is correct in UK tax law.

Calculation correctness4.89 /5

Whether the steps and final figures are right.

Tax year correctness4.91 /5

Whether rates, thresholds and rules match the relevant tax year.

Citation correctness4.91 /5

Whether the cited sources support the specific claims made.

Coverage

What the bank covers, and what is still being built

The bank covers 20 UK tax domains and seven use-case families. The current 176 questions are almost entirely Family A (information retrieval). Families B, C and D have a handful of questions each. Families E, F and G have not yet been authored. The roadmap below shows each family's progress toward its planned share of the 1,000-question target.

Information retrieval

Find a fact in the law and report it correctly.

  • A1Look up rates and thresholds
  • A2Interpret a specific authority
  • A3Find the procedural mechanics
167/200Established

Applied advisory

Apply the law to a client's facts and reach a position.

  • B1Apply law to one issue with full facts
  • B2Advise on a complex situation with rich context
  • B3Handle a sparse-context query
4/300In progress

Quantitative analysis

Compute figures and compare numerically.

  • C1Compute a liability or relief
  • C2Compare scenarios
  • C3Find a tipping point
4/200In progress

Strategic advisory

Plan ahead and recommend a course of action.

  • D1Optimise the tax outcome
  • D2Recommend a structure
  • D3Balance trade-offs
1/100In progress

Defensive review

Find what is wrong or exposed in a position.

  • E1Spot risks
  • E2Review a return for errors
  • E3Assess defensibility against HMRC
0/100Planned

HMRC engagement

Deal with HMRC procedure, enquiries and penalties.

  • F1Respond to an enquiry
  • F2Plan a voluntary disclosure
  • F3Analyse penalties
0/50Planned

Specialist & boundary

Handle niche, changing and genuinely unsettled areas.

  • G1Cross-border
  • G2Anti-avoidance
  • G3Recently changed law
  • G4Genuinely unsettled areas
0/50Planned
DomainABCDEFGNowTarget
Anti-Avoidance22·····450
Capital Allowances1·1····250
Capital Gains Tax23·1····2450
Charity Tax2······250
Compliance9······950
Construction Industry1······150
Corporation Tax1311····1550
Digital Services Tax1······150
Employment Tax11······1150
Income Tax27·1····2850
Inheritance Tax33··1···3450
International Tax101·····1150
National Insurance7······750
Partnerships1······150
Pensions1······150
Property Taxation2······250
Share Schemes1······150
Stamp Duty Land Tax4······450
Trusts & Estates1······150
Value Added Tax17······1750
Total · 20 domains167441···1761000
1-4 5-14 15+
A RetrievalB AppliedC QuantitativeD StrategicE DefensiveF HMRCG Specialist
Progression

Every run is scored and recorded

We re-run the benchmark continuously. The share of answers that are correct or partially correct climbed from 89% to 98% across recent runs, while the question bank grew from 97 to 176 questions.

Correct or partially correct, per run. Bank grew from 97 to 176 questions over the period. The dip is a real run, shown as recorded.

Methodology

How we score

The Golden answer is the correct, expert-written answer to each question: the position in UK tax law, the figures for the relevant tax year, and the supporting citations to primary or secondary legislation or HMRC material. Every GAIN Tax answer is scored against its Golden answer.

Our tax specialist writes every question and Golden answer, then validates every result and signs off the report. Each answer is scored on eight dimensions against a fixed rubric. Three are critical: a failure on law, calculation or tax year caps the result regardless of the rest.

Legal correctnesscritical25 %
Calculation correctnesscritical25 %
Tax year correctnesscritical25 %
Citation correctness25 %
Source relevancesupporting
Completenesssupporting
Fabricated / unsupported claimsupporting
Practical usefulnesssupporting
  1. 01Questions and Golden answers are written by our tax specialist and fixed before the run.
  2. 02Each answer is scored against its Golden answer on the eight criteria above.
  3. 03Our tax specialist validates every result and signs off the final report.

What this number is, and is not

This is our own question bank, currently 176 questions and moving toward 1,000 across 20 domains. It measures performance on these questions, at this version. We are open that a different set of questions can produce a different result.

Our aim is not a marketing claim. It is to give tax professionals an honest view of where these tools are strong, where they are weak, and what the risks are. GAIN Tax is a research tool, and outputs should be verified before they reach a client.

We welcome any tax professional who wants to join the benchmark and test the tool objectively.

About Us

About GAIN Tax

GAIN Tax is AI-powered UK tax research software for accountants, tax advisers and tax lawyers. It answers UK tax questions with citations to primary and secondary legislation, HMRC manuals and HMRC guidance, so every answer can be checked against its source.

The product is built on retrieval-augmented generation over UK tax law, kept up to date, with a Client Hub that holds client profiles, documents and history so answers reflect the full client context. GAIN Tax is a research tool that makes tax work faster. It does not replace professional judgment, and we recommend verifying every output.

GAIN Tax is led by its two founding brothers: a finance executive and ACCA-qualified chartered accountant serving as CEO, and an AI and technology leader serving as CTO. The company builds tools that bring rigour and transparency to the use of AI in tax.

This benchmark is part of that aim: to give tax professionals an objective, evidence-based view of where AI is strong, where it is weak, and what the risks are.

Get in touch →