.png)
A commercial invoice lists 12,500 kg. The bill of lading lists 12,850. The packing list lists 12,600. All three documents are legible. OCR can even extract the information from them without difficulty. Yet something is wrong.
Which of these figures should we focus on?
That is where the real challenge with logistics documents lies. Simply reading the information isn't enough. You also need to understand what it refers to, compare it with other information in the file, and identify when there are discrepancies.
OCR makes it possible to read a document. Document-based AI goes a step further: it can identify the document’s content, extract useful data, and make that data actionable.
But in a shipping operation, documents don’t stand alone. An invoice must match a packing list; a bill of lading must match the shipment details; and a customs declaration must be supported by the relevant documentation.
The difference between OCR and document-based AI, therefore, lies not only in the ability to read a document. It lies in what can be done with the information once it has been extracted.
And that is where document reconciliation, consistency checks, and—more broadly—a case-centric approach to AI begin.
OCR, which stands for Optical Character Recognition, converts text in an image or PDF into usable digital data.
This technology is essential for digitizing paper or scanned documents. In logistics, for example, it can be used to recognize:
But recognizing a piece of information does not mean understanding its role.
If a document contains “12,500 kg,” OCR can read this value perfectly well. It does not necessarily know whether this is the gross weight, the net weight, or some other piece of information. Nor does it know whether this figure matches the one listed on the bill of lading or the packing list.
What OCR Doesn't Do
OCR primarily processes the content of a single document. On its own, it cannot:
This technology therefore remains a fundamental building block of document processing. But it answers a simple question:
"What does this document say?"
In logistics, the following question is often more important:
"Does this information align with the rest of the file?"
Document-based AI applied to transportation and logistics goes beyond OCR. It doesn't just read a document; it seeks to extract useful information from it and make it directly actionable.
In particular, it can:
A shipping order arrives via email, sometimes accompanied by a PDF or other attachments.
Using an automated document processing approach, the system can identify the order, extract the necessary information from the email and PDF, and send the structured data to the TMS to pre-fill the shipping order.
The operator no longer re-enters the same data. Instead, the operator verifies the information and takes action when necessary.
This is a significant change: the document is no longer just digitized. It becomes a data source that can be directly utilized by the business process.
.png)
An international shipment rarely relies on a single document. A single shipment may be described in a commercial invoice, a packing list, a bill of lading, a customs declaration, or a freight invoice.
Each one contains a portion of the information. The problem arises when these pieces of information do not match.
Case Study:
OCR can read all three values. Document-based AI can extract and structure them.
But one question remains: Why is this data different?
This is where document reconciliation. It involves comparing information from multiple documents to detect discrepancies, verify consistency, and identify anomalies.
The system no longer simply tries to determine what each document contains. It tries to understand whether the documents tell the same story.
The difference between data extraction and information management is most evident in their uses.
A shipping order arrives via email with one or more attachments.
Documentary AI identifies documents, extracts relevant information, organizes it, and sends it to the TMS; it can also pre-fill a freight forwarder’s shipping record in the TMS.
The result: less data re-entry and a lower risk of errors.
All three documents contain the same information: references, quantities, weights, and values.
AI can extract them and then compare them to detect discrepancies.
Result: An inconsistency is identified before it causes a delay, a billing error, or a documentation issue.
Before a declaration is submitted, several checks can be performed on the file: to verify the presence of the required documents, the consistency of the information, the source data, the classification, or other applicable requirements.
The goal is not to replace the reporter, but to alert them to any discrepancies and items that need to be verified before submission.
Result: The filer is working on a file that has already been reviewed, with the discrepancies identified.
A shipping invoice can be compared to the quote or the agreed-upon pricing terms.
AI extracts data from the invoice, matches it to reference information, and flags any discrepancies.
As a result, billing errors can be detected before they are approved.
These four examples illustrate the process: extracting data is just the first step. The real value emerges when that data can be compared, verified, and used in the business process.
A logistics operation relies on several documents that sometimes arrive at different times: shipping orders, invoices, packing lists, bills of lading, customs documents, and updated versions.
Treating them separately results in some context being lost. To detect inconsistencies or verify the compliance of a transaction, the AI must be able to link documents related to the same shipment and cross-reference their information.
The reasoning then no longer focuses solely on:
"What is in this document?"
but on:
"What can we infer from all the documents related to this operation?"
It is this shift in scale that makes it possible to move beyond data extraction and apply controls that are truly useful to logistics operations.
A document-centric, agentic approach is not simply a matter of adding an AI agent to a document management tool.
The agent has access to the context of the documents related to a single transaction. The agent can then analyze the information, apply business rules, identify anomalies, and propose a course of action.
For example, after detecting a discrepancy between an invoice, a packing list, and a bill of lading, it can:
Oversight remains essential. In critical document-processing workflows, the goal is not to let AI operate unchecked, but to combine automation, business rules, and human intervention when necessary.
Auditability is just as important: a decision or alert must be traceable to the data and documents that supported it.
This brings us from:
extract → structure
To:
understand → monitor → make recommendations → act under supervision.
It is this ability to reason within a business-specific documentary context that distinguishes a document-centric agentic approach from simple extraction automation.
OCR reads documents. Document-based AI extracts data from them. But the real value emerges when it becomes possible to link, verify, and leverage that data within the business process.
In logistics, the challenge is therefore no longer just to read documents, but to ensure the reliability of the information they contain.