Technologies

What is a data annotation?

Data annotation involves associating structured information with raw data to enable an artificial intelligence model to identify, classify, or interpret it.

Data can be annotated in various ways depending on the model and use case: an image can be associated with a category, a text with an intent or an entity, and a document with a type or specific fields.

Annotation is therefore an important step in the development of many artificial intelligence systems. It provides models with examples from which they learn to recognize information and replicate certain behaviors.

But when it comes to processing business documents, annotating data does not necessarily mean understanding a document.

This distinction becomes essential when a process is based on multiple documents that describe the same operation.

What is the purpose of data annotation in artificial intelligence?

An artificial intelligence model does not inherently possess a business-specific understanding of the data presented to it.

Annotation involves providing information that allows a piece of data to be associated with a meaning or a category.

Let's take a simple example.

A system must automatically identify the different types of documents received by a company. Documents are assigned to categories such as:

  • invoice;
  • packing list;
  • Bill of Lading;
  • certificate of origin;
  • customs document;
  • shipping document.

These annotated examples are then used to train or improve a document classification model.

The same principle applies to data extraction. A document is annotated to indicate where certain pieces of information are located: invoice number, date, amount, currency, weight, quantity, reference number, or the name of a party.

Annotation helps the model learn what to look for and how to interpret certain information.

But it doesn't necessarily answer a key question:

What does this information mean when compared to the other information in the file?

Data annotation and document annotation: What's the difference?

In a documentary process, a document is generally not an isolated source.

An international shipment may include a commercial invoice, a packing list, a bill of lading, a customs declaration, and various certificates.

Each of these documents contains some of the information needed to process the transaction.

An annotation allows a system to correctly identify:

  • the type of document;
  • the important fields;
  • the associated values;
  • References;
  • quantities;
  • the dates;
  • the amounts.

However, even a correctly identified piece of data may still be inconsistent with the rest of the file.

Let's consider a shipment for which the documents state:

Document Weight
Invoice 10,000 kg
Packing List 9,800 kg
B/L 10,200 kg

The three pieces of information may have been extracted correctly. However, there is an anomaly in the file.

So the problem is no longer about extracting the data. It lies in the relationship between the data.

This is where document automation takes on a whole new character.

From Annotation to Document Comprehension

In a document management system, multiple processing stages may occur sequentially.

1. Identify

The system determines the type of document received. Is it an invoice, a packing list, a B/L, or a customs document?

2. Extract

The relevant information is identified in the document.

3. Organize

The information is converted into data that can be used by business systems.

4. Bring closer

The data is compared with the data in the other documents in the file.

5. Check

The information is cross-checked against business rules, reference data, or data already available in the company's systems.

6. Decide

When an anomaly occurs, the system can flag the problem and provide guidance on the necessary action.

This progression makes it possible to distinguish between intelligence applied to a document and intelligence applied to a file.

This is precisely the shift that document reconciliation entails: moving from processing documents in isolation to verifying the relationships between the information they contain.

Why can data that has been correctly extracted still be incorrect when viewed in context?

This is one of the main challenges of document automation. Information can be perfectly legible and correctly extracted, yet still be unusable for the rest of the process.

Let's take an invoice that lists 10,000 kg, while the packing list lists 9,800. An extraction system may find both documents to be correct.

It isn't necessarily an OCR error.

There isn't necessarily an extraction error.

But there is an inconsistency in the documentation.

This difference is important.

OCR enables reading. IDP enables extraction and structuring.

Document reconciliation allows you to verify the consistency of a file.

Document automation should therefore not merely aim to produce more data. It should aim to produce data that is reliable enough to be used in business processes.

Annotation, OCR, IDP, and document reconciliation: What are the differences?

These technologies and approaches address different problems.

Technology / Approach Primary Function
Data Annotation Preprocess data to enable a model to learn or be evaluated
OCR Recognize the text in a document
IDP Identify, extract, and organize information from a document
Document reconciliation Compare information from multiple sources and identify inconsistencies
Document Automation Moving data and documents through a business process

This distinction is important because these bricks are not interchangeable.

A system can certainly extract data without knowing whether that data is consistent with other data in a different document.

Do you want to be more productive?

Book a demo
Book a demo

Data Annotation Applied to Logistics Documents

In logistics and international trade, companies handle a wide variety of documents from numerous sources on a daily basis.

A single operation involves:

  • commercial invoice;
  • packing list;
  • Bill of Lading;
  • Air Waybill;
  • CMR;
  • certificate of origin;
  • customs declaration;
  • shipping invoice;
  • regulatory documents.

Annotation helps train systems to identify these documents or extract relevant information from them. But the operational value becomes apparent when this information can then be consolidated, compared, and verified.

You can compare the weight listed on an invoice with that on a packing list.

A quantity can be reconciled with another document.

A reference can be cross-checked against the operational file.

An HS code can be verified against the description of the goods.

A shipping invoice can be compared to the applicable rate schedule.

A required document may be identified as missing.

The question then becomes:

Is the file complete, consistent, and usable?

Why Document Reconciliation Is Becoming Essential

International document flows do not consist of a single document that is processed once and for all. Documents arrive at different times, and the information changes as they are processed.

That is why effective automation must be able to take into account the historical context of the case.

This approach makes it possible to shift from a document-processing approach to one focused on file consistency, capable of incorporating new information and verifying its consistency with existing data.

The file serves as a natural tool for document reconciliation: it allows for maintaining a consolidated and contextualized view of the operation and for continuing to perform checks as the file evolves.

Manual vs. Automated Annotation: What Are the Differences?

Annotation can be done manually, automatically, or using a hybrid approach.

Manual annotation

An operator assigns labels or identifies the necessary information. This approach allows for precise control but becomes costly as volumes increase.

Automated Annotation

Some models automatically suggest or assign labels. This makes it possible to process more data, but it requires control and validation mechanisms.

Hybrid Approach

In many professional settings, the goal is not to completely eliminate human involvement.

Rather, it involves automating repetitive tasks and having an operator step in when the system encounters uncertain information or a situation that requires a business decision.

This " human-in-the-loop " approach is particularly important when the data has operational, financial, or regulatory implications.

What are the use cases for data annotation?

Annotation plays a role in many areas of artificial intelligence. It is used to train models capable of:

  • classify documents;
  • recognize entities in a text;
  • identify objects in an image;
  • extract information;
  • categorize content;
  • detect certain anomalies.

In the document workflows of international trade, these capabilities support the processing of commercial, logistics, customs, and financial documents. However, requirements vary depending on the industry.

For freight forwarders

A freight forwarder must manage document flows from numerous parties and systems. The challenge, therefore, is not only to extract information from each document, but also to consolidate it into a usable file and detect inconsistencies before they become operational problems.

For customs declarants

In customs operations, the completeness and consistency of information are essential. A missing document, inconsistent information, or incorrect data can result in a request for clarification, a delay, a risk of noncompliance, or even a fine.

Document automation can therefore be used prior to filing to verify the available information and identify areas requiring attention.

For Trade Finance

Trade finance applications are also based on a set of documents that must meet specific requirements.

The challenge, then, is to verify not only that the expected documents are present, but also that the information they contain is consistent: amounts, dates, references, currencies, parties, goods, or shipping documents.

Document automation can help identify discrepancies before they require further verification or delay the processing of the case.

Is data annotation enough to automate a document-processing workflow?

No. That's probably the most important distinction to keep in mind. Annotation allows us to label data and contribute to the training of AI models.

But a real documentary process usually involves more than that:

identify → extract → organize → reconcile → verify → decide → transmit.

The difficulty increases even further when multiple documents describe the same transaction and the documents arrive gradually.

In this context, a system's performance should not be evaluated solely on its ability to correctly extract information.

We should also ask ourselves:

  • Is the data consistent with other sources?
  • Is the file complete?
  • Is a business rule being followed?
  • Was an anomaly detected?
  • Should we require human validation?
  • Can the data be transmitted to the business system?

It is this logic that shifts the focus from data automation to process automation.

Toward a New Generation of Document Intelligence

Advances in information technology are gradually following this trend.

OCR has made it possible to convert documents into usable text.

The IDP added classification, extraction, and structuring.

Document reconciliation approaches provide a key capability: comparing information from multiple sources and detecting discrepancies.

Document reliability goes even further by seeking to make data consistent, complete, and usable enough to support business processes. This approach is one of the key aspects of Docloop’s positioning in the field of Trade Document Intelligence.

So the challenge is no longer simply to ask the AI:

"What does this document say?"

But let's take it step by step:

"What does this information mean in the context of the case? Is it consistent with other sources, and what action should be taken?"

It is this shift in perspective that paves the way for truly operational document automation.

FAQs
No items found.