Three lines

Uber

Developers

Transcription

Transcription is the process of converting audible speech from audio or video into accurate, time-synced text. This work type enables you to analyze, search, and document spoken content with professional-level accuracy.

Transcription is essential for any task that requires analyzing or searching the content of spoken words.

Ideal for: Call center analytics, content searchability, medical documentation, meeting analysis, accessibility, and any scenario requiring conversion of speech to text.

When to Use

Use Transcription when your input data is audio or video containing spoken content, and your primary goal is to convert that audible speech into accurate, time-synced text.

Simple Speech-to-Text Tasks (Content Conversion)

Ideal for converting raw voice recordings into a text format for analysis, documentation, or searchability. The output is typically plain text, potentially with basic punctuation and speaker labels.

Input Type Question Example Purpose / Requirement
Audio “Convert this recorded customer service call into text.” Call Center Analytics: Analyzing sentiment, identifying common issues, or ensuring quality control.
Video “Transcribe the dialogue spoken in this lecture video.” Content Searchability: Making video and audio files searchable and accessible (e.g., YouTube subtitles/captions).
Audio “Transcribe the patient’s dictation from this medical recording.” Medical Documentation: Converting spoken medical notes into digital records.

Time-Sensitive and Rich Transcription Tasks

Use this when you need not only the words but also precise timing or additional context about the speaker or environment.

Input Type Question Example Purpose / Requirement
Audio “Transcribe the interview, ensuring each word is tagged with its start and end time (timestamps).” Alignment and Synchronization: Precisely linking the text to the audio for subtitling, forced alignment, or detailed analysis.
Multi-Speaker Audio “Transcribe the meeting and label which speaker said each sentence (Speaker 1, Speaker 2, etc.).” Speaker Diarization: Analyzing multi-person interactions, like identifying who spoke how often in a meeting.
Video “Transcribe the narration, noting any non-speech events like ‘[laughter]’ or ‘[applause]’.” Accessibility and Context: Providing a complete text representation of the audible experience for those who cannot hear it.

Key Features

  • Timestamp Precision - Get exact start and end times for each word or phrase
  • Speaker Identification - Distinguish between multiple speakers in conversations
  • Non-Speech Annotations - Capture ambient sounds like laughter, applause, or music
  • Custom Formatting - Apply specific punctuation, capitalization, or formatting rules

When to Use Other Work Types

Supported Data Types

  • Audio - MP3, WAV, FLAC, and other audio formats

Getting Started

  1. Define your transcription requirements (timestamps, speakers, formatting)
  2. Specify any domain-specific terminology or formatting rules
  3. Prepare your audio files
  4. Submit batches via API
  5. Track progress and download text transcripts

How to Submit Batches

Indirect Demand (CSV File) - Best for large-scale projects with thousands of audio files.

📖 View API Documentation →

Direct Demand (Inline Data) - Best for dynamic tasks or real-time processing.

📖 View API Documentation →

Best Practices

  1. Audio Quality - Ensure clear audio with minimal background noise for best results
  2. Speaker Context - Provide context about number of speakers and roles when possible
  3. Domain Terminology - Share industry-specific terms or jargon for accurate transcription
  4. Formatting Guidelines - Specify punctuation, capitalization, and timestamp requirements upfront
  5. Quality Control - Review sample transcripts to ensure accuracy and consistency

Uber

Developers
© 2026 Uber Technologies Inc.