PDF ExtractionAI

Why "99% accuracy" means nothing in PDF data extraction — and what to ask instead

6 min read  ·  Nextraxion

Every document extraction tool claims 99% accuracy. It's the headline figure on every pricing page, in every sales deck, in every comparison table. And it is almost completely meaningless.

Not because the vendors are lying. But because "accuracy" as a single number tells you almost nothing about what will happen when you upload your actual documents from your actual suppliers to their actual product.

Here's what the number doesn't tell you — and what you should be asking instead.

The accuracy percentage problem

When an extraction vendor says "99% accuracy," they typically mean: across a test set of documents, 99% of extracted values matched the expected output. What they don't tell you is what documents were in that test set. Clean, digital PDFs generated by accounting software? Or scanned invoices from a 1990s fax machine, photographed at an angle and compressed to 72dpi?

They also don't tell you which fields were included in the calculation. If "invoice number" extracts correctly 100% of the time but "line item unit price" fails one in three times, you can still claim 99% accuracy overall — while delivering data that's completely unusable for your accounts payable workflow.

A single accuracy percentage hides the fields that matter to you behind the fields that are easy to get right.

The fields that fail most often in real-world extraction are exactly the ones that matter most: line items, tax breakdowns, cost centre codes, PO references, and anything that appears in a table or non-standard layout. These are also the fields where manual correction takes the longest — so a tool that fails on them is saving you almost nothing.

Why traditional OCR gets this wrong

OCR — Optical Character Recognition — is the original PDF extraction technology and it's still the foundation under many modern tools. OCR reads text from documents character by character. Feed it a scanned invoice and it converts the image to a string of characters. What it cannot do is understand what those characters mean.

OCR can read "£1,234.56" perfectly. It has no idea whether that number is a subtotal, a VAT amount, a line item price, or a deposit. It captures the text. Context is your problem.

This is why raw OCR tools require you to map fields manually — "the total is always in position X, Y on the page." That works until a supplier updates their invoice template. Then it breaks silently: the extraction still runs, but the values land in the wrong fields. You won't know until someone notices a payment error weeks later.

Template-based extraction — better, but brittle

The next generation of tools added templates on top of OCR. Define a rule: "the invoice total is the number immediately to the right of the text 'Amount Due'." More reliable — until a supplier redesigns their invoice, moves the label, uses "Balance Payable" instead, or starts using a two-column layout.

Template-based tools require ongoing maintenance. Every new supplier is a new template. Every supplier invoice redesign is a broken template. For a business processing invoices from 40 different suppliers, that's 40 templates to build and an unknown number to fix whenever something changes.

How extraction technology has evolved
OCR
Limited

Reads characters. No understanding of meaning or context.

Template-based
Brittle

Rule-based field mapping. Breaks whenever a layout changes.

AI / LLM
Resilient

Semantic understanding. Adapts to any layout, any supplier.

How AI extraction is different

AI-powered extraction — the approach Nextraxion uses — applies large language models that understand documents the way a human reader would. They don't look for text in a fixed position. They understand that "Amount Due," "Balance Payable," and "Total Outstanding" are the same concept regardless of where they appear on the page, and regardless of which supplier sent the invoice.

No templates. No position mapping. No retraining when a supplier changes their layout. You define the fields you want — invoice number, vendor name, line items, VAT, due date — and the model works out where they are on every document it processes.

This is not just faster to set up. It's fundamentally more resilient. New supplier? Upload the invoice. Supplier redesigned their template? Upload the invoice. The model adapts without you touching anything.

The metric that actually matters: field-level confidence

Here's what to ask any extraction vendor instead of "what's your accuracy?": do you provide a confidence score on every extracted field, individually?

Nextraxion does. Every field in every extracted document gets its own confidence score — a signal that tells you how certain the model is about that specific value. High confidence on vendor name and invoice number. Lower confidence on a handwritten annotation in the margin. You see this at the field level, not as an average across the document.

This matters because it changes your workflow fundamentally. Instead of either trusting everything (and missing errors) or reviewing everything (and defeating the purpose of automation), you review only the fields that need it. High-confidence fields go straight through. Low-confidence fields go to a human reviewer. The validation queue handles the exceptions — automatically, without you deciding case by case.

Field-level confidence turns extraction from a binary "did it work?" into a managed workflow where humans focus only on the decisions that actually need them.

Extracted fields — Harrington Supplies Invoice #4471Nextraxion
Vendor nameHarrington Supplies Ltd
✓ 99%
Invoice numberINV-4471
✓ 98%
Invoice date03 Jun 2026
✓ 97%
Subtotal£8,340.00
✓ 95%
VAT (20%)£1,668.00
✓ 94%
Line item 4 — unit price£147.50
⚠ 61%
PO referencePO-2026-0183
⚠ 58%

2 fields flagged for review · 5 fields auto-approved

What this means in practice

A tool that claims 99% accuracy with no confidence scoring will extract 990 fields correctly out of 1,000 — and give you no indication which 10 are wrong. You either review everything (slow) or trust everything (risky).

A tool with field-level confidence might extract 985 fields at high confidence and flag 15 for review. You spend two minutes on the 15. The 985 go straight to your accounting system. Your effective throughput is higher, your error rate is lower, and your team only touches the decisions that need human judgment.

That's the difference between a headline number and a workflow that actually works.

Questions to ask before you choose an extraction tool

Before committing to any PDF data extraction platform, ask these:

  • What documents were in your accuracy test set — and can I test on mine?
  • Do you provide confidence scores per field or per document?
  • What happens when confidence is low — does it fail silently or flag for review?
  • How do you handle suppliers who update their invoice layout?
  • Do I need to build templates, or does the model adapt automatically?

The answers will tell you more than any percentage.

See field-level confidence in action

Upload a PDF and see exactly which fields extract at high confidence and which need a second look — before you commit to anything.

Try free extraction →