Column Tax 发布 TaxCalcBench,评测前沿模型能否完成报税计算
TaxCalcBench: Can AI file your taxes?
Column Tax 发布 TaxCalcBench,用 50 个报税案例评测 38 个前沿模型能否完成报税的计算步骤,答案与预期申报表逐行比对。2025 版(v2)比 2024 版更难:输入改为真实的 PDF 文档(W-2、1099 等)而非结构化数据,案例同时包含州税和联邦税,税务与财务情形也更复杂。
Form 1040-AI
Column Tax—Benchmark of Frontier Models on Tax Calculation
Model Use Only—Do not write or hallucinate in this space.
For the tax year Jan. 1–Dec. 31, 2025, or other benchmark edition beginning below
Benchmark identification number TCB | 25 | v2
Branch main
Returns tested 50
State CA IL NY VA
Models tested 38
Model Submission Campaign
Check here if you, or your model, want to be benchmarked. Email [email protected]. Checking a box below will not change your score.
Ranking Status
Check only one box.
Web Search
At any time during the benchmark, did the model (a) use a web search tool, or (b) otherwise look up tax forms and instructions online? (See instructions.) Show filers who answered
Providers
(see instructions)
If more than four providers, see below and check here
Results
Attach model outputs here. Also attach Schedule M for any line.
If your model is not listed, see instructions.
Click any line to open Sch. M.
Form 1040-AI (2025) Page 2
Refund
Amount You Owe
Third Party Designee
Do you want to run this benchmark on your own model? See instructions.
Yes. Complete below. No
Phone no.uv sync --all-extras
License (PIN)MIT
Sign Here
Joint return? See instructions. Keep a copy for your records.
Under penalties of perjury, I declare that I have examined this leaderboard and accompanying schedules and statements, and to the best of my knowledge and belief, they are true, correct, and complete. Declaration of model (other than taxpayer) is based on all tokens of which the model has any knowledge.
Your signatureColumn Tax
Date10/01/2026
Your occupationTax engine builders
If the model refused to prepare a return, enter the reason here (see inst.)
Model's signature. If a joint return, both must sign.The Models
Date10/01/2026
Model's occupationTax preparer (in training)
Paid Preparer Use Only
Preparer's nameColumn Tax
Preparer's signatureColumn Tax
2025
Instructions for Form 1040-AI
TaxCalcBench Leaderboard
What's New
Tax Year 2025 (v2 edition). The 2025 edition is harder than the 2024 edition in three ways: inputs are realistic PDF documents (W-2s, 1099s, and so on) instead of structured data, cases include state returns as well as federal returns, and the cases cover much more complex tax and financial situations.
Tax Year 2024 (v1 edition). Federal-only returns for relatively simple situations, with inputs provided as structured JSON. Each case was run four times and the scores were averaged (pass@1). Use the tax year boxes at the top of the form to switch editions.
General Instructions
What is TaxCalcBench?
TaxCalcBench measures whether frontier AI models can do the calculation step of tax preparation: given everything about a taxpayer, produce the completed return. Each test case pairs a taxpayer's inputs with the expected, correctly computed return, and the model's answer is compared to it line by line.
Tax calculation has traditionally been done by hand-built, deterministic tax engines that encode tens of thousands of pages of rules. This benchmark asks whether a model can do the same job on its own.
Who must file
Model providers who want their model on this form can email [email protected]. Anyone can also run the open-source harness from the repository.
Line Instructions
Column (a)—Correct returns (strict). The share of test cases where every evaluated line exactly matches the expected return. This is the number that matters: a tax return has to be completely correct to be filed.
Column (b)—Correct returns (lenient). The share of cases where every evaluated line is within ±$5 of the expected value. Many small misses come from computing tax with bracket math instead of the official tax tables.
Column (c)—Correct (by line). The average percentage of evaluated lines that exactly match. A single early mistake can cascade through the rest of a return, so this is usually much higher than column (a).
Column (d)—Correct (by line, lenient). The same as column (c), counting lines within ±$5 as correct.
Column (e)—Cost per return. Average API cost to produce one return. A blank means cost wasn't available; blanks are never treated as $0. 2025 edition only.
Column (f)—Time per return. Average generation time for one return. 2025 edition only.
Thinking level. Each model is tested across its supported reasoning levels, and each column reports the best setting for that column. Two numbers on the same line can come from different settings.
Web search. Lines marked Web search had a web search tool available, so the model could look up current forms and instructions.
Partial coverage. Lines marked with a case count, such as 40/50, are scored only on the cases that completed. The TaxCalcBench README explains each of these runs.
Specific Instructions
Model-specific notes are attached to each model's Schedule M. Click a line on page 1 to open it.
Paperwork Reduction Act Notice
We ask for the information on this form to find out whether a model can do your taxes. Models are not required to provide the information requested unless they want to be on the leaderboard. The average time burden for completing this form varies by model; see column (f). If you have suggestions for making this form simpler, we would be happy to hear from you at [email protected].
Results synced from the TaxCalcBench repository on 10/01/2026.
来源:Hacker News · AI · taxcalcbench.ai