Blog
BPO4 min27 July 2026

AP outsourcing accuracy benchmarks: what best-in-class actually looks like

Under 0.8% error rate, under $1 per invoice, ~6,900 invoices per FTE per year — the IOFM numbers that separate a top-tier AP operation from an average one, and what drives the gap.

"We're accurate" is not a benchmark. IOFM's accounts-payable shared-services research puts real numbers behind what separates an average operation from a best-in-class one — and the gap is large enough that it changes what a client should expect to ask for in an outsourcing contract.

The numbers

MetricAverage performerBest-in-class
Cost per invoice$10-22 (manual-heavy)Under $1 (AI-driven)
Error rateMaterially higherUnder 0.8%
Invoices processed per FTE per year~4,200~6,900

The cost gap alone — a 10-22x range depending on where an operation sits — is the headline most outsourcing conversations focus on. The error-rate and throughput numbers are arguably more important, because they compound: a lower error rate means fewer costly corrections downstream, and higher throughput per person means an operation can absorb client growth without a proportional headcount increase.

Where the gap actually comes from

It is tempting to attribute the difference to "better software" in the abstract. The more precise answer is that best-in-class operations have eliminated the specific failure modes that manual and template-based processing cannot avoid:

Manual keying introduces errors that automated extraction does not

Every hand-typed figure is a chance for a transposition error, and error rates compound with volume — a clerk keying hundreds of invoices a day makes more mistakes on invoice #400 than invoice #40, simply from fatigue. Automated extraction does not degrade with volume.

Template maintenance caps throughput

A template-based system requires someone to build and maintain a layout map per supplier. That work does not show up in "cost per invoice processed" — it shows up as a permanent, growing maintenance burden that silently caps how many invoices per FTE a team can realistically sustain, because part of every FTE's time goes to template upkeep rather than invoice processing.

Sampling-based checks miss what systematic checks catch

An average operation checks a sample of invoices for arithmetic and duplicate errors, because checking every invoice by hand does not scale. A best-in-class operation running automated checks on every document catches the errors that would otherwise surface downstream — at month-end, or worse, at a client audit.

What "under 0.8% error rate" actually requires

Getting error rates that low is not primarily about a smarter model — it is about systematic checks running on every document, every time: arithmetic consistency (do line items sum correctly), tax-rate coherence (does the charged rate match what is valid for that supply type), and duplicate detection against full history rather than the current batch alone. None of these checks are individually sophisticated; the achievement is running all of them, on everything, without exception.

This is precisely how DOXALIO's pipeline is structured: every extracted figure is source-cited to its page, arithmetic and tax coherence run on every line of every document, and duplicate detection compares each new invoice against a client's full history, not just the current batch — including mutated invoice numbers a human reviewer would not remember from three months back. Bank-detail changes on known suppliers surface automatically, before the invoice reaches a payment queue.

What this means for evaluating an AP outsourcing partner

Ask for the actual numbers, not the adjective. A partner should be able to state their error rate, their cost per invoice, and their throughput per FTE — and explain what changed operationally to reach those figures, not just claim "AI-powered" as if the label were the achievement.

Related reading

FAQ

Is under $1 per invoice realistic for smaller operations, or only at scale?

The floor gets easier to approach as fixed setup costs disappear — template-based systems needed volume to amortize the per-supplier configuration cost, which favored scale. Understanding-based document AI has no comparable per-supplier setup cost, which changes the economics for smaller operations too.

Does a 0.8% error rate mean 0.8% of invoices are paid incorrectly?

No — it typically refers to extraction and coding accuracy at the field level, with flagged low-confidence items routed to human review before anything posts. The design principle behind the low error rate is that nothing posts without either high machine confidence or human approval.

How is throughput per FTE actually measured?

As invoices fully processed (extracted, checked, coded, and either auto-approved within policy or queued for a fast human decision) per employee per year — a number that rewards eliminating queue time and re-work, not just typing speed.

Ready to try DOXALIO?

Free trial. No credit card required.

Get started for free
AP outsourcing accuracy benchmarks: what best-in-class actually looks like — DOXALIO Blog