Guide · Operational Practice

Document Automation:

Document automation is one of the use cases with the clearest ROI in business. But not every document is equally easy to automate. Which document types can be processed reliably today, which stubbornly stay manual, and how a pilot is set up in practice.

A practical guide for accounting, procurement and logistics teams at DACH companies — no marketing promises, just realistic recognition rates.

Short answer

Document automation works reliably today for standard invoices, delivery notes, purchase order confirmations and simple freight documents, with typical recognition rates of 90–95 percent. What stays difficult is handwritten documents, poor-quality faxes, contracts with individual wording and rare languages — here, 60–80 percent is realistic. The key isn't the highest possible automation rate, but clean escalation logic: uncertain documents go to a human with a stated reason, instead of being processed blindly.

1. What works reliably today

Four document types can be processed automatically in practice with recognition rates between 90 and 95 percent — provided volume is high enough and the data structure is stable:

Standard invoices from known suppliers

Recurring suppliers with a stable layout are the classic case. Supplier recognition, field extraction (date, amount, VAT, purchase order reference), hand-off to the ERP — all automatic. Manual handling only for exceptions (e.g., a quantity mismatch against the order).

Delivery notes with a structured layout

Delivery notes in the standard B2B format are usually well structured. Line-item recognition, quantity matching against the order, automatic goods-receipt posting — all achievable in production.

Purchase order confirmations

Incoming purchase order confirmations can be matched against the original order. Discrepancies (date, quantity, price) are automatically flagged and escalated — with no manual step in between in the standard case.

Simple freight documents and forwarding

Delivery papers, freight-forwarder documents, simple customs paperwork — these mostly follow standardized structures and automate well, especially combined with a track-and-trace system.

2. What stubbornly stays manual

Just as important as what works is what doesn't. Four document types remain difficult in practice — recognition rates of 60–80 percent, a higher escalation rate, and more ongoing maintenance:

  • Handwritten documents: notes, corrections, filled-in forms. Even modern handwriting recognition rarely achieves a reliable 90 percent. Realistic: 60–75 percent, with clear escalation logic for uncertain spots.
  • Poor-quality faxed documents: resolution, shadows, skew, toner streaks — all of these kill reliable recognition. Preprocessing helps here (deskewing, threshold adjustment), but no model can fix arbitrarily bad input.
  • Contracts and legal documents with individual wording: unlike structured documents, contracts have no fixed fields. Automation is possible here (clause classification, risk flagging), but expecting "full processing" is unrealistic.
  • Documents in rare languages or mixed languages: standard models are optimized for a handful of languages. Documents in rare languages or with mixed languages (e.g., English + Chinese + German in one waybill) often reach only 60–70 percent automatic processing.

3. Escalation logic is the key

The most common marketing mistake in document automation: promising "100 percent automation." In practice, the real goal is different — clean escalation logic that recognizes uncertain documents and hands them to a human with a stated reason, instead of waving them through blindly.

A well-built solution has three escalation tiers:

L1

High confidence → fully automatic processing

The document is read, validated, and handed off to the ERP. No human step. Typically 70–85 percent of documents.

L2

Medium confidence → processing with pre-filled data

The document goes to a human, but with fields already pre-filled. The person reviews and confirms instead of entering everything from scratch. Typically 10–20 percent.

L3

Low confidence → fully manual handling

The document is too difficult for automatic processing. It goes to a human along with the reason for the escalation ("handwritten correction unclear," "unknown format"). Typically 5–15 percent.

4. Integration: usually half the effort

Document recognition is one side of the solution. The other — and in practice the more effort-intensive one — is integration with the ERP, accounting, and the document management system (DMS). Without that integration, even the best recognition is worthless: the data ends up sitting in a spreadsheet nobody maintains.

Typical target systems at DACH companies: SAP, BMD, DATEV, Microsoft Dynamics 365, Oracle, Sage 100/X3. Most offer REST APIs or structured import formats. Before choosing a model, this question needs an answer: "How does the recognized document get into the target system — and who maintains that interface when the target system changes?"

Rule of thumb: if you estimate implementation effort only for recognition and don't price in the integration, you're building a demo, not a production system.

5. What a document automation pilot looks like

A realistic document automation pilot looks like this in practice:

Weeks 1–2

Preparation & baseline

Define the document type (e.g., invoices). Measure the baseline: time per transaction, error rate, data-entry time. Document the interface to the ERP/accounting system.

Weeks 3–5

Build & ERP connection

Train or configure the model on specific suppliers. Build the interface. Define escalation logic with confidence thresholds.

Weeks 6–8

Parallel operation

The solution runs on the real document flow. Human processing runs in parallel as a control. Edge cases are collected and folded into the escalation rules.

Week 9+

Expansion or stop

Based on the pilot data: expand to further document types, or deliberately stop. Both options are legitimate — what matters is the data behind the decision.

When document automation pays off — and when it doesn't

Makes sense when
  • • At least 50–100 structured documents per day
  • • Recurring suppliers or document types
  • • A target system with a documented interface (ERP, DATEV, BMD)
  • • A named operational owner for escalation cases
  • • Willingness to start with a 70–85 percent automation rate
Doesn't make sense when
  • • Under 30 documents per day (effort outweighs benefit)
  • • Document types change significantly every week
  • • The target system has no open interfaces
  • • Expecting "100 percent automatic, no escalation"
  • • Nobody owns the escalation cases

Frequently asked questions about document automation

Which document types can be processed automatically best today?
Standard invoices from known suppliers, delivery notes with a clearly structured layout, purchase order confirmations, simple freight documents. Common traits: a recurring layout, legible text, and a clear field position. Typical recognition rates for these document types run at 90–95 percent automatic processing; the rest is escalated.
Which documents remain difficult in practice?
Handwritten documents, poor-quality faxed documents, contracts with individual wording, documents in rare languages, documents with handwritten corrections, and documents with attachments that themselves contain multi-page structures. Realistic recognition rates here run at 60–80 percent — the rest needs human handling.
Does document automation pay off at low volume already?
Below 50 documents a day, a fully automated solution rarely pays off. Between 50 and 200 a day, ROI is typically reached within 6–9 months. From 200 documents a day, 4–8 months is realistic. For very small volumes, a partially automated solution (pre-filling plus a human in the loop) can make more sense than full automation.
How do you handle edge cases?
Edge cases aren't avoided — they're escalated. A well-built solution recognizes on its own when a document can't be processed reliably, and hands it to a human reviewer, with a stated reason for the escalation. This escalation rate should be measured transparently. A rate under 15 percent is typical; under 5 percent is exceptional.
How does document automation integrate with existing ERP/accounting systems?
Through the existing interfaces of the ERP/accounting system. The usual suspects: SAP, BMD, DATEV, Microsoft Dynamics, Oracle, Sage. Most systems offer REST APIs or structured import formats (XML, CSV, EDI). Integration is usually half the implementation effort — and the reason the interface analysis has to happen BEFORE model selection.

Document automation for your document flow?

In a free initial call, we assess your specific document flow: which document types can realistically be automated today, which stay manual, and where the economic leverage lies.

Request a free initial call