Meet @O.

Online Workshop on 14 October: Survive The AI EraOnline Workshop: Survive The AI EraRegister
Free 30-min AI audit

Benchmark report · September 2026

How accurately does AI read accounting documents?

We ran four model setups over 4’801 receipts and invoices and 400 bank statement files. Here is what each one gets right, what it costs, where it fails — and what that means for your review and for your clients’ data.

Key findings

An incoming invoice is only useful for bookkeeping once its date, number, total and VAT amount have been captured correctly. Ogment reads those details from your documents so that your team checks them instead of typing them in. Which AI models do that reading, and how well, is a decision we make on evidence. This report is that evidence: every setup we ran in full, measured on the same documents against known correct answers.

One distinction runs through all of it: reading a document is not booking it. The tests measure whether information was extracted correctly. They do not measure VAT treatment, the choice of account or a bank reconciliation. Those remain your professional decisions.

of receipt and invoice totals read correctly by the most accurate setup
99.24%
of tax amounts on real receipts, best setup: the weakest result on receipts and invoices
93.18%
what the most accurate setup costs compared with our production route, on receipts and invoices
11×
of transaction dates and balances on bank statements read correctly by the most accurate setup, for transactions matched to the reference
99.99%
  • The best setup is very accurate on the fields that matter most. Kimi K3 read 99.89% of document types, 99.80% of document numbers, 99.24% of totals and 99.15% of dates correctly. Our production route is within two thirds of a percentage point of it on each of those fields.
  • Real receipts are where the errors are. On synthetic documents the three strongest setups (the production route, Kimi K3 and the partner model) read each of the five key fields correctly at least 99.39% of the time. On real shop and restaurant receipts, they read the tax amount correctly between 89.32% and 93.18% of the time.
  • More accuracy costs a lot more. Kimi K3 costs USD 39.49 per 1’000 documents against USD 3.47 for the production route: about 11 times as much, for a gain of 0.13 to 0.63 percentage points on each of the five exact-match fields.
  • A model small enough for one server in Switzerland was not good enough. Qwen3-VL-30B is cheap and its weights are open, but it read only 85.19% of tax amounts and 93.04% of document types correctly, and completed 326 of 400 bank statement files.
  • Many errors are blanks, not wrong values. In the production run, 31 of 52 subtotal errors, 18 of 38 tax errors and 13 of 40 date errors were fields left empty. An empty field is visible to a reviewer; a wrong number is not.

What we tested

A setup is the pair of models that does the work: one reads each page, one extracts the accounting fields from that reading. Our production route uses a different model for each step, each with a second model as a fallback. The three challengers use one model for both steps and no fallback, so each result shows what that model does on its own. A model’s weights are the files that make up the trained model; when they are published (“open”), anyone can run the model on their own servers.

The four setups

SetupReads the pagesExtracts the fieldsModel weights
Production routeGPT-5 mini; fallback Claude Haiku 4.5GPT-5.6 Luna; fallback Claude Opus 4.6Closed
Kimi K3Kimi K3Kimi K3Open; about 1.4 TB
Partner modelProprietary partner modelThe same modelClosed
Qwen3-VL-30BQwen3-VL-30BQwen3-VL-30BOpen (Apache 2.0); 62.2 GB
The production route is the same pair of models and instructions that processed customer documents at the time of the tests, run here in a test harness on the test documents. The partner model is a proprietary model from a partner, reached through its own interface; we do not name it, and it has no public price.

Three more models did not get a full run. Two, Qwen3.5-27B and Qwen2.5-VL-72B, could not work with our document-processing pipeline at all. Mistral Small 3.2 completed 95 documents of a 100-document trial and read 68.10% of their tax amounts correctly, which did not justify a full run.

Receipts and invoices

The first test set is 4’801 documents from three public research datasets: 1’987 real receipts from shops and restaurants in Indonesia and Malaysia, and 2’814 synthetic (computer-generated) invoices and receipts. Every document comes with reference answers. For each field, a result is the share of reference answers that the setup extracted correctly, averaged over the datasets that have a reference answer for that field, with each dataset counted equally. The real receipts are foreign, so “tax amount” means the sales or service tax printed on them, not Swiss VAT.

Share of fields read correctly, by setup

What was measuredProduction routeKimi K3Partner modelQwen3-VL-30B
Document type: receipt or invoice99.76%99.89%99.73%93.04%
Document number99.39%99.80%99.59%99.26%
Total amount98.61%99.24%98.92%94.65%
Document date98.53%99.15%98.71%98.28%
Tax amount shown on the document95.96%96.51%94.53%85.19%
Supplier name (matching score)95.81%96.10%96.04%93.92%
Supplier address (matching score)94.97%95.78%95.37%93.20%
Line-item descriptions (matching score)93.70%94.73%93.33%92.61%
Line items with their amounts (matching score)96.11%96.46%96.04%94.58%
Documents processed, after retries100.00%100.00%100.00%99.04%
4’801 documents, September 2026; the best value in each row is in bold, here and in the tables that follow. The first five rows are exact matches. The supplier and line-item rows are matching scores that penalise both missing and superfluous content, so they are not comparable with the rows above. Document numbers have a reference answer in the synthetic dataset only. Qwen3-VL-30B was given one retry only; the other three were retried until every document completed.

Three of the four setups are close. Kimi K3 is best or equal best on every row, the production route and the partner model trade places behind it, and the three are within two points of each other on every exact field. Qwen3-VL-30B is a different story: competitive on dates and document numbers, clearly behind on tax, totals and document type.

A result applies to one field at a time. 99.24% for the total means that about 99 in 100 totals were extracted correctly. It does not mean that 99.24% of documents were entirely correct, and it does not describe the share of correct accounting entries.

Real receipts against synthetic documents

What was measuredProduction routeKimi K3Partner modelQwen3-VL-30B
Total amount: real receipts (CORD)98.15%99.28%98.05%86.83%
Total amount: real receipts (SROIE)97.67%98.48%98.78%97.87%
Total amount: synthetic documents100.00%99.96%99.93%99.25%
Tax amount: real receipts (CORD)92.05%93.18%89.32%–
Tax amount: synthetic documents99.87%99.83%99.75%–
Totals and tax amounts by dataset. CORD and SROIE are real receipts; the tax amount has a reference answer on 440 of the CORD receipts and on none of the SROIE ones. “–” marks a slice that was not published for that setup.

Wrong values per 1’000, most accurate setup

  • Document type

    Real receipts 1.5

    Synthetic documents 0.4

  • Total amount

    Real receipts 11.2

    Synthetic documents 0.4

  • Document date

    Real receipts 14.2

    Synthetic documents 2.8

  • Tax amount

    Real receipts 68.2

    Synthetic documents 1.7

Kimi K3: how many of every 1’000 reference answers were read wrongly or left empty, on real receipts and on synthetic documents.

The pattern is the same for the three strongest setups: synthetic documents are nearly solved, real receipts are not. That matters for how to read any accuracy claim, including ours. A benchmark made only of clean, generated invoices would have shown each of the three strongest setups above 99% on each of the five exact-match fields.

Bank statements

The second test set is 400 bank statement files: 200 synthetic statements, modelled on Indian business bank statements, each as a digital PDF and as a scanned copy. A statement is a harder task than a receipt: several pages, many transactions, and every row has to be found before its fields can be right.

Files processed to a complete result

What was measuredProduction routeKimi K3Partner modelQwen3-VL-30B
Completed on the first pass360 of 400365 of 400337 of 400279 of 400
Completed after one retry382 of 400384 of 400372 of 400326 of 400
Share of all 400 files95.50%96.00%93.00%81.50%
Share of the 388 files the test setup accepted98.45%98.97%95.88%84.02%
Results in the required structure100.00%99.48%99.48%87.63%
One retry was allowed for files that failed on the first pass. Twelve of the 400 files have five pages where the test expects six, and were set aside before any model saw them; the fourth row leaves only those twelve out.

No setup completed every file. Beyond those twelve files, the production route failed on 6 files, Kimi K3 on 4, the partner model on 16 and Qwen3-VL-30B on 62.

Statement details read correctly, by setup

What was measuredProduction routeKimi K3Partner modelQwen3-VL-30B
Transaction date99.98%99.99%99.86%98.21%
Balance after each transaction99.88%99.99%99.95%98.12%
Transaction amount99.69%99.59%99.79%99.24%
Debit or credit98.60%99.98%99.83%93.82%
Transaction text (matching score)98.71%99.53%98.36%92.79%
Transaction rows found (row score)97.49%97.85%96.02%87.41%
Closing balance95.50%95.75%93.00%81.50%
Account number95.50%95.25%93.00%81.50%
Statement period95.50%95.25%93.00%52.25%
The first five rows are counted on transactions that were found and matched to the reference. The last four are counted over all 400 files, so every file that failed counts as wrong, including the twelve set aside before processing (3 points for every setup); the row score also counts transactions that were missed or added. On the files it completed, the production route read the closing balance, account number and period correctly every time.

On transactions that were found, the three strongest setups read dates, amounts and balances correctly at least 99.59% of the time. The differences are elsewhere. The production route confuses debits and credits more often than Kimi K3 or the partner model: 98.60% against 99.98% and 99.83%. Qwen3-VL-30B has the lowest row score, and got the statement period right on barely half of the files.

Scanned copies against digital originals

What was measuredProduction routeKimi K3Partner modelQwen3-VL-30B
Files completed+1.00 pts−1.00 pts−3.00 pts−14.00 pts
Transaction rows found (row score)+0.94 pts−1.02 pts−2.90 pts−14.04 pts
Scanned minus digital, in percentage points, on the 200 paired statements. The production route’s small gain on scans is noise from retries, not evidence that scans are easier.

Scans cost the partner model about three points and Qwen3-VL-30B fourteen. If your clients send scans and photographs rather than bank exports, this is the table to look at.

Cost and speed

Accuracy is one side of the decision. The other is what a setup costs to run and how long a document takes.

Cost per 1’000 receipts and invoices

  • Qwen3-VL-30B

    USD 1.36

  • Production route

    USD 3.47

  • Kimi K3

    USD 39.49

US dollars at the providers’ published prices in September 2026, for the runs reported here. Qwen3-VL-30B’s figure is observed usage and leaves out 40 failed calls. The partner model has no public price and is not shown.

Receipts and invoices: cost and median time

What was measuredProduction routeKimi K3Partner modelQwen3-VL-30B
Cost per 1’000 documentsUSD 3.47USD 39.49No public priceUSD 1.36
Median time per document200 s40 s11 s12 s
The lowest value in each row is in bold. US dollars at the providers’ published prices in September 2026, for the results reported here. Qwen3-VL-30B’s cost is observed usage only; its 40 failed page-reading calls are not included. Median time runs from sending a document to receiving its result; each run was measured under a different load, so the times show an order of magnitude, not a controlled speed test, and the production route’s is inflated by a queue in front of the page-reading step.

Bank statements: cost and median time

What was measuredProduction routeKimi K3Partner modelQwen3-VL-30B
Cost per 100 filesUSD 9.33USD 49.35No public priceUSD 11.10
Median time per file452 s252 s201 s734 s
The lowest value in each row is in bold. Costs are the usage the providers reported for the whole run, failed attempts included; calls with incomplete usage data are not counted, so these are lower bounds. A statement file has five or six pages.

The most accurate setup is also the most expensive by a wide margin: Kimi K3 costs 11.4 times the production route on receipts and invoices and 5.3 times on bank statements, on the usage the providers reported. For bank statements we looked at what that buys — 0.50 points more files completed, 0.36 points on the row score, 1.38 points on debit and credit — and kept the production route.

Qwen3-VL-30B is the cheapest setup on receipts and one of the two fastest. On bank statements it is neither: it is the slowest setup and costs more than the production route.

Where it goes wrong

Averages hide what an accountant most needs to know: which mistakes to expect. We went through the errors field by field. These are the patterns, with counts from the production-route run unless another setup is named.

  • Tax amounts on real receipts

    The weakest result on receipts and invoices. Of the 440 real receipts with a reference tax amount, the production route got 405 right, Kimi K3 410 and the partner model 393. The receipts that go wrong typically have a missing tax line or an ambiguous zero tax; on others a digit is misread or the amount is derived wrongly. Nearly half of the production route’s tax errors, 18 of 38, are fields left empty rather than wrong amounts.

  • Thousands separators and scale

    Indonesian receipts write thirty-six thousand as 36.000, and an earlier version read some of these as 36. The instructions now tell the model to read such separators consistently with the arithmetic printed on the receipt. That, with the other changes of the same round, raised correct totals on those receipts from 83.13% to 97.63%. The same ambiguity exists wherever conventions differ: 1’234.50 and 1.234,50 are the same amount.

  • Dates

    The production route got 40 of 3’801 dates wrong, and left thirteen of those empty. 23 of the 40 are on scanned receipts, where the usual cause is a misread digit or a compact date that can be read in two ways; 17 are on synthetic documents.

  • Document numbers

    An earlier version often returned the printed label together with the number: “Order #INV-2024-089” instead of “INV-2024-089”. That caused 19 of its 34 errors and is fixed. The 15 errors that remain are three empty fields and dropped or substituted characters.

  • Receipts mistaken for invoices

    In an earlier version, 736 of the 987 receipts in one dataset were classified as invoices. After we rewrote how the two are told apart, 981 of those 987 are recognised as receipts. The chart below shows the change for each group of documents.

  • Blank fields

    When the production route fails on a field, it often returns nothing rather than a wrong value: 31 of 52 subtotal errors, 18 of 38 tax errors, 13 of 40 date errors and 10 of 41 total errors were empty fields. For a reviewer this is the better kind of error, because an empty field is visible and a wrong number is not.

  • Runs that do not finish

    Every model sometimes fails to return a usable result: a page-reading call times out, or the answer does not have the required structure. On receipts and invoices the production route needed 3 retries, Kimi K3 had 6 failures on the first pass and the partner model 12, and all three finished every document. Qwen3-VL-30B failed on 128 documents on the first pass and still had 46 unfinished after a retry.

  • Failed files and direction on bank statements

    On bank statements the largest loss is a file that fails entirely. Among the files that complete, the production route’s main weakness is direction: it assigned debit or credit wrongly on 1.40% of the transactions it found, against 0.02% for Kimi K3.

Documents given their right type, before and after the fix

  • Real receipts (SROIE)

    Before 24.52%

    After 99.39%

  • Synthetic receipts

    Before 83.28%

    After 100.00%

  • Real receipts (CORD)

    Before 98.80%

    After 100.00%

  • Synthetic invoices

    Before 99.95%

    After 99.84%

The same 4’801 documents, production route. One group got slightly worse: 3 of 1’857 synthetic invoices are now misclassified, against 1 before.

A second look at the arithmetic

An invoice checks itself: subtotal plus tax equals the total, and the line items add up to the subtotal. After the comparison above we built that check, together with a second extraction pass, and reran the production route with both on 10 September, on a larger set of 5’301 documents that adds 500 synthetic card-payment slips (card-terminal receipts, not Swiss payment slips).

of 5’301 documents whose extracted amounts did not add up
709
subtotals changed by the second pass
232
tax amounts changed by the second pass
43
of card-payment slip totals read correctly
99.80%

This run uses a different set of documents, so its results cannot be merged into the tables above. What it shows is how often the arithmetic disagrees, and that a second pass does change values: most often the subtotal (232 documents), then the supplier address (62), the currency (60), the document number (44) and the tax amount (43). Whether each change is a correction is a separate measurement.

Could the AI run in Switzerland?

Fiduciaries ask whether the AI models themselves could run in Switzerland, and not only the storage. We looked at what that would take for the two setups whose model weights are open.

  • Qwen3-VL-30B fits on one server. Its weights are 62.2 GB and fit on a single 96 GB graphics card. A server with that card lists at CHF 2’632 per month at one Swiss provider and at under CHF 2’993 at another, before operations, backup and a second server for availability. But this is the setup that read 85.19% of tax amounts correctly and completed 326 of 400 statement files.
  • Kimi K3 needs a cluster. Its weights are about 1.4 TB. Sixteen graphics cards in a Swiss data centre come to about CHF 24’800 per month at list price, for a capacity we estimate at five to ten simultaneous users. That is not a basis for serving a product.
  • The production route and the partner model cannot be self-hosted. Their weights are not published.

So today, among the setups we tested, the choice is between a model we could host in Switzerland and a model that is accurate enough, and we chose accuracy. Your documents are stored in Switzerland; the AI processing takes place in the European Union, on the terms described below.

Prices: Public list prices in Swiss francs, checked on 4 September 2026 for the Kimi K3 cluster and on 9 September 2026 for the single server. The capacity figures are estimates, not measurements.

How we measured

  1. Documents with known answers. We used public research datasets in which each document comes with reference answers: the date, total or transactions actually on it. No client documents were used. Of the receipts and invoices, 1’987 are real receipts and 2’814 are synthetic; all bank statement files are synthetic.
  2. Processing. Each setup processed every document the way the product does: each page was read, and the accounting information was extracted into defined fields such as date, document number, total and tax amount.
  3. Comparison with the reference. Each extracted value was compared with the reference answer. Common differences in date format did not count as errors: 31.03.2024 and 2024-03-31 are the same date. Amounts had to match to the cent; document numbers had to match apart from spaces, punctuation and upper or lower case.
  4. Counting. For each field, we counted how many reference answers were extracted correctly. Documents without a reference answer for a field were left out for that field; where a reference answer existed but nothing was extracted, it counted as an error. Because 59% of the receipts and invoices are synthetic, each score was calculated within each dataset first and the datasets were then given equal weight, so that real receipts count as much as synthetic documents.

For illustration, not a test case: If an invoice shows a total of CHF 108.10, an extracted 108.10 counts as correct, while 108.01 or 1’081.00 counts as an error.

On bank statements, an extracted row was paired with a reference row only when the two largely agreed across date, amount, balance, direction (debit or credit) and description, and the pairs had to follow the order of the statement. Date, amount and balance were then checked on the paired rows. A row read too differently to be paired does not count as a wrong amount: it lowers the row score instead, as a missed row and, if it was extracted, as an extra one.

The datasets are CORD v2 (1’000 real receipts from Indonesia), ICDAR 2019 SROIE (987 real scanned receipts from Malaysia), Invoice OCR Synthetic (2’814 synthetic invoices and receipts) and Indian Bank Statements (400 synthetic files). The first three are published under the CC BY 4.0 licence and the fourth under Apache 2.0, and we used fixed versions of each.

What this means for your review

  • The results describe the test documents. The real receipts are shop and restaurant receipts from Indonesia and Malaysia, not Swiss documents. Your suppliers’ layouts, languages and scan quality can differ from what we tested.
  • Part of the test data is synthetic. Document numbers were measured on synthetic documents only, and all bank statements are synthetic rather than real Swiss statements. As the results above show, synthetic documents are easier than real ones.
  • One accurate field does not make a whole document correct. Each field is measured on its own.
  • Scans are harder. On bank statements, results on scanned copies were up to three points lower than on the digital originals for the three strongest setups.
  • Test results are not a promise. They describe these runs in September 2026, not every document, and the setup that processes your documents can change when another one does better in these tests.
  • Professional judgment stays with you. The tests did not cover VAT treatment, account assignment, reconciliation or the correctness of accounting entries.

In practice: check extracted tax amounts and totals — particularly on receipts — before you rely on them, look twice at empty fields, and review statement data as you would any other input to a reconciliation.

Where your client data is processed

Your clients’ documents are confidential. Three separate questions determine what happens to them, and we answer each one separately.

  • Where are the documents stored?

    In Switzerland. The documents you upload to Ogment are stored in Switzerland.

  • Where does the AI processing take place?

    In the European Union only. When a document is read and its information extracted, that AI processing takes place exclusively in the EU.

  • What do the AI providers keep?

    Not your document content. Under our Zero Data Retention policy, our AI providers do not retain the document content they receive or the responses they generate once processing is complete. Separately, they do not use this content to train AI models.

Zero Data Retention applies to the AI providers. It does not mean that Ogment deletes your documents: they stay stored in your Ogment workspace in Switzerland, so your team can keep working with them. And “not used for training” and “not retained” are two different commitments — we make both.

Teo BorschbergCEO, Ogment

Book your free 30-min AI audit

  • Where your team’s hours go, mandate by mandate
  • Which workflows an AI agent can take over first
  • A clear AI ROI estimate for your firm
Company size