Can AI Close the Books? What the Evidence Shows for Accounting Work
Can AI close the books? The evidence so far points to two answers at once. On clearly defined accounting tasks, current AI models now beat junior accountants on speed, accuracy and cost. On the full job of a month-end close, no model is reliable without a person checking the work.
Last reviewed: October 2026. We review this page every month and list changes at the bottom.
Most of the evidence comes from Mercor, a company that builds AI benchmarks and licenses expert-built datasets. It created the APEX-Accounting benchmark together with Ramp, so treat the numbers as vendor-run research that is unusually open about its limits.
What AI does well
In October 2026, Mercor published a study in which 12 licensed CPAs, averaging about five and a half years of experience, worked through four month-end close scenarios adapted from its benchmark. Each one required finding the right figures in a company's files, doing the math and returning a table of results. Mercor compared them with AI models given the same tasks.
The accountants met about 37% of the grading criteria on average. Individual attempts ranged from 0% to about 90%, and most took 30 to 180 minutes. The best model in the study scored full marks on all 20 of its attempts, each finished in under 10 minutes. Mercor also estimated the cost per grading criterion met at about $0.21 for that model, against about $10.35 for the accountants, based on the US median accountant wage.
Mercor was careful about what this shows. The tasks were designed to trip up models, and they reward close reading of files and strict instruction following, which is where AI is strongest. The accountants had no colleagues to ask and no background on the company, and the study did not test client communication or knowing which question to ask.
Where it still fails
The study used simplified tasks. The full APEX-Accounting benchmark has 160 tasks across 10 company environments, each frozen at month-end close, with 2,186 grading criteria. Mercor says more than 40 professionals built the tasks, with a median of 11 years of experience and half from Big Four firms, and that every model ran each task eight times.
The results are far less flattering. Mercor reports that 58% of the tasks were never fully solved by any model in any of the eight attempts. The top mean scores on the public leaderboard sit around 61 to 62% of criteria met, as of early October 2026. Mercor also reports that models can solve a task once but rarely solve it every time, that spending more per run did not buy reliability, and that over half of the top model's errors were multi-step accounting reasoning mistakes.
Why both results are true
The two results measure different things. A clear, bounded task with the right files in front of you is something AI now does well. A real close involves incomplete records, company-specific context and many dependent steps. Mercor notes that requirements in its tasks compounded, so a single overlooked number could produce a very low score.
How to read an accounting AI headline
- Who ran it? Vendor-run studies can still be useful, but check who built the tasks and who sells what.
- How big was the sample? Twelve accountants is a small group.
- What counts as a pass? Meeting some grading criteria is not the same as getting the whole task right.
- Was it repeated? A model that succeeds once is different from one that succeeds every time.
What to check before you choose a tool
This section is our reading, not a finding from the research. If you are comparing Month-End Close, Bookkeeping or Audit & Compliance tools, the evidence suggests asking:
- Which steps does the tool automate, and which does a person approve? Defined tasks such as matching, coding and reconciling are the safer place to start.
- Does it show its sources and workings so a reviewer can check a figure quickly?
- Can you run it more than once on a closed period from your own books and compare it with the answer you already know?
- How does it handle missing or conflicting records: does it flag them or guess?
- What does it cost per run, or per close, at your volume?
Update log
- October 2026: First version. Added Mercor's October 1 human-baseline study and the APEX-Accounting leaderboard (about 61 to 62% top mean score, 58% of tasks unsolved). Next review: November 2026.
Sources:
Verified By
GuideToReviews Team
Need a custom AI tool recommendation?
Our AI Assistant tracks the latest mergers and updates to recommend the best tools for your specific workflow.