Three lines

Uber

Developers

Document Annotation

Document Annotation is the process of annotating specific regions, fields, or text boxes within scanned documents, PDFs, or form images for structured data extraction. This task is focused on teaching models to visually understand the layout and hierarchy of documents.

Document Annotation bridges the gap between Computer Vision (layout recognition) and Natural Language Processing (text extraction), often in preparation for Optical Character Recognition (OCR) or structured field extraction.

Ideal for: Form processing, invoice automation, identity verification, document structure analysis, OCR training, and any scenario requiring structured data extraction from document images.

When to Use

Use Document Annotation when your input is a scanned document, a PDF, or a form image, and you need to annotate specific regions, fields, or text boxes within that document for structured data extraction.

Layout and Region Labeling Tasks

Ideal when you need to define the structural components of a document image, such as detecting tables, headers, or form fields before extracting the content.

Input Type Question Example Purpose / Requirement
Form Image “Draw a bounding box around the Table Area, the Signature Field, and the Total Amount Field.” Form Processing: Teaching a model to isolate and identify specific zones for automated processing.
Scanned PDF “Mark the regions corresponding to the Page Header, the main Body Text, and the Footer.” Document Structure Analysis: Understanding the logical flow and layout for better indexing and navigation.
Invoice Image “Label the bounding box that contains the Invoice ID and the one that contains the Vendor Name.” Key-Value Pair Extraction: Locating the visual positions of critical data points for extraction.

OCR Correction and Text Markup Tasks

Use this when you need to specifically label text regions to improve OCR accuracy or to associate extracted text with its visual location.

Input Type Question Example Purpose / Requirement
Document Image “Draw a box tightly around every word and label it with the transcribed text.” OCR Training/Tuning: Providing ground truth for training or improving Optical Character Recognition models.
Check Image “Mark the box around the Written Amount and the box around the Numerical Amount.” Financial Automation: Isolating specific text fields that require high-precision recognition.
ID Card Scan “Draw precise polygons around the text of the Name field and the Date of Birth field.” Identity Verification: Extracting personal information from standardized document templates.

Key Distinction

Document Annotation is unique because the input is visual (an image of text), but the task is to produce structured text data (what the field is and where it is located). It bridges the gap between Computer Vision (layout recognition) and Natural Language Processing (text extraction).

Common Document Types

  • Forms - Applications, questionnaires, surveys
  • Invoices - Bills, receipts, purchase orders
  • Contracts - Legal documents, agreements
  • ID Documents - Passports, driver’s licenses, identity cards
  • Medical Records - Patient forms, prescriptions, lab reports
  • Financial Documents - Checks, bank statements, tax forms

When to Use Other Work Types

Supported Data Types

  • PDF - Native and scanned PDF documents

Getting Started

  1. Define your document types and field requirements
  2. Specify field locations and bounding box precision
  3. Prepare your PDF documents
  4. Submit batches via API
  5. Track progress and download structured field annotations

How to Submit Batches

Indirect Demand (CSV File) - Best for large-scale projects with thousands of documents.

📖 View API Documentation →

Direct Demand (Inline Data) - Best for dynamic tasks or real-time processing.

📖 View API Documentation →

Best Practices

  1. Clear Field Definitions - Define precise criteria for each document field or region
  2. Template Consistency - Document variations in layout across document types
  3. OCR Requirements - Specify text extraction accuracy requirements
  4. Boundary Precision - Define how tight bounding boxes should be around text
  5. Quality Control - Review sample annotations to ensure field identification accuracy

Uber

Developers
© 2026 Uber Technologies Inc.