Skip to content

Claude beats accountants: 100 percent in closing test

A Mercor test shows Claude solving simplified monthly closing tasks flawlessly in 20 of 20 attempts, while twelve experienced auditors clearly lag behind. In the full 160-task APEX-Accounting benchmark, even the best model meets only about 62 percent of the criteria. Client conversations and years of experience are excluded, as routine bookkeeping is now faster and cheaper with AI.

By Brian Beckmann · 3 October 2026 · 3 min

A robotic arm with a Claude logo stacks file folders on a scale that tilts to its side, while a person across from it at a calculator has only stacked a few folders

The HR platform Mercor tested AI models against twelve licensed auditors on real monthly closings. In four simplified closing scenarios, Claude completed all 20 test runs flawlessly, each in under ten minutes, while the professionals, with an average of 5.5 years of experience, performed markedly worse.

On the full, 160-task exam, even the strongest model reaches only about 62 percent of the grading criteria.

Claude masters simplified closing tasks

For the study, Mercor, the US HR platform specialized in AI evaluations, had twelve licensed auditors compete against an AI model from Anthropic. The task: four realistic monthly-closing scenarios, where the needed figures first have to be dug out of a company's working files before being calculated and turned into a results table.

According to Mercor, Claude completed all 20 test runs flawlessly, each in under ten minutes. The professionals, averaging 5.5 years of experience, performed markedly worse; their scores ranged from zero to about 90 percent.

Mercor's task authors had expected averages of 30 percent for juniors and 55 percent for mid-level professionals going in — the test group's actual performance came in below that. For companies already offloading repetitive closing work to tools like Copilot in Excel, this confirms a trend: structured, rule-bound grunt work is now handled more reliably by AI than by many junior professionals.

Full benchmark exposes clear limits

The second, far tougher exam is called APEX-Accounting and tests models on 160 tasks across ten simulated companies with 2,186 grading criteria in total; each task requires systems to independently work through an average of 7.5 source documents, with no file list or prescribed method provided.

The tasks and grading rubrics were authored by professionals from major auditing firms, including Deloitte, PwC, EY, and KPMG. Each model ran the exam eight times to reveal fluctuations in reliability.

On the public APEX-Accounting leaderboard, Claude Opus 5.5 leads with 61.8 percent of criteria met, closely trailed by Fable 5.1 at 61.0 percent and GPT-6 Astra at 57.9 percent; other tested models such as Gemini 3.1 Pro, Grok 4.5, and Qwen3.5-397B rank further behind.

The numbers show that even the strongest model misses well over a third of the criteria — fully automated bookkeeping remains out of reach for the current generation of AI.

Progress against earlier models stands out

The leap looks bigger across time: eighteen months ago, the best AI models of that era, according to Mercor's own, independently unverified figures, still trailed the historical benchmark of roughly 37 percent for average accountants.

Today the leading model reaches nearly 62 percent on the full exam, and scores a perfect result on the simplified closing scenarios. Deliberately left out are tasks that go beyond pure number-crunching: conversations with clients, check-ins with colleagues, and the understanding of a single company's quirks built up over years.

That is precisely where human judgment remains irreplaceable for now — the benchmark targets exactly what AI already does best: tracking down evidence and following instructions precisely.

What matters next is whether auditing firms use the gap between simple routine work and complex closings to shift junior staff toward client contact and plausibility checks rather than cutting positions outright. Mercor says it will keep updating the APEX-Accounting leaderboard as new model versions arrive — the next data point should show whether the gap between AI and human professionals keeps narrowing on the harder tasks.

Sources

  1. Mercor: Human baselines for benchmarks — AI now outperforms junior accountants
  2. Mercor APEX-Accounting Leaderboard

Common questions

Your AI update for the work week

Once a week, the most important AI news – plus one practical tip to try right away. No spam, unsubscribe anytime.

/ki-praxis/2026-10/claude-schlaegt-buchhalter-abschlusstest /en/blog/2026-10/claude-beats-accountants-closing-test