Three lines

Uber

Developers

Object Detection

Object Detection is the process of identifying, counting, and precisely locating one or more objects within an image or video frame. The core output is a bounding box (or similar shape) and a class label for every instance of an object of interest.

Object Detection is a combination of Image Classification (telling you what the object is) and Object Localization (telling you where the object is).

Ideal for: Autonomous driving, surveillance, retail inventory, medical diagnosis, sports analytics, and any scenario requiring precise object location and identification.

When to Use

Use Object Detection when you need to identify, count, and precisely locate objects within an image or video frame.

Single-Object Localization Tasks

Ideal when you need to confirm the presence of an object and pinpoint its location, even if there’s only one instance. This task is crucial for applications that must interact with physical items.

Input Type Question Example Purpose
Image “Draw a box around the traffic light and label its state (Red/Yellow/Green).” Autonomous Driving: Identifying and understanding critical, single-instance road elements.
Video Frame “Locate the primary subject’s face in this frame and label the person’s emotion (Happy/Neutral/Sad).” Facial Recognition/Pose Estimation: Tracking specific features or body parts.
Medical Scan “Draw a bounding box around any detected tumor and classify its type.” Medical Diagnosis: Pinpointing abnormalities in X-rays or MRI scans.

Multi-Object Identification and Counting Tasks

Use this when the image/video contains multiple instances of the same or different objects that all need to be cataloged and located. This is the primary strength of Object Detection over simple Image Classification.

Input Type Question Example Purpose
Image “Draw a box around every car, truck, and pedestrian.” Surveillance/Traffic Monitoring: Tracking multiple entities in a scene for safety or flow analysis.
Video “Track the location of every player and the football throughout this clip.” Sports Analytics: Real-time tracking of game entities for tactical analysis.
Image “Locate and label every product box visible on the shelf.” Retail/Inventory Management: Automated shelf auditing to check stock levels.
Aerial Image “Draw a box around and count all solar panels in this satellite image.” Environmental/Asset Monitoring: Large-scale feature counting and condition checks.

Key Difference from Classification

The fundamental distinction is the output format:

  • Classification gives a single label for the whole image (e.g., “The image contains a car.”)
  • Object Detection gives multiple bounding boxes and multiple labels for objects within the image (e.g., “Car at coordinates [x1, y1, x2, y2], Pedestrian at [x3, y3, x4, y4].”)

When to Use Other Work Types

Supported Data Types

  • Image - Detect and locate objects in static images
  • Video - Detect and track objects across video frames

Getting Started

  1. Define your object categories and detection requirements
  2. Prepare your image or video data
  3. Submit batches via API
  4. Track progress and download results with bounding box coordinates

How to Submit Batches

Indirect Demand (CSV File) - Best for large-scale projects with thousands of images or videos.

📖 View API Documentation →

Direct Demand (Inline Data) - Best for dynamic tasks or real-time processing.

📖 View API Documentation →

Best Practices

  1. Clear Object Definitions - Define precise criteria for what constitutes each object class
  2. Bounding Box Guidelines - Specify how tight boxes should fit around objects
  3. Overlapping Objects - Provide guidance for handling partially occluded objects
  4. Edge Cases - Document how to handle objects at image boundaries
  5. Quality Control - Review sample outputs to ensure consistency across annotators

Uber

Developers
© 2026 Uber Technologies Inc.