Transcription
Transcription is the process of converting audible speech from audio or video into accurate, time-synced text. This work type enables you to analyze, search, and document spoken content with professional-level accuracy.
Transcription is essential for any task that requires analyzing or searching the content of spoken words.
Ideal for: Call center analytics, content searchability, medical documentation, meeting analysis, accessibility, and any scenario requiring conversion of speech to text.
¶ When to Use
Use Transcription when your input data is audio or video containing spoken content, and your primary goal is to convert that audible speech into accurate, time-synced text.
¶ Simple Speech-to-Text Tasks (Content Conversion)
Ideal for converting raw voice recordings into a text format for analysis, documentation, or searchability. The output is typically plain text, potentially with basic punctuation and speaker labels.
| Input Type | Question Example | Purpose / Requirement |
|---|---|---|
| Audio | “Convert this recorded customer service call into text.” | Call Center Analytics: Analyzing sentiment, identifying common issues, or ensuring quality control. |
| Video | “Transcribe the dialogue spoken in this lecture video.” | Content Searchability: Making video and audio files searchable and accessible (e.g., YouTube subtitles/captions). |
| Audio | “Transcribe the patient’s dictation from this medical recording.” | Medical Documentation: Converting spoken medical notes into digital records. |
¶ Time-Sensitive and Rich Transcription Tasks
Use this when you need not only the words but also precise timing or additional context about the speaker or environment.
| Input Type | Question Example | Purpose / Requirement |
|---|---|---|
| Audio | “Transcribe the interview, ensuring each word is tagged with its start and end time (timestamps).” | Alignment and Synchronization: Precisely linking the text to the audio for subtitling, forced alignment, or detailed analysis. |
| Multi-Speaker Audio | “Transcribe the meeting and label which speaker said each sentence (Speaker 1, Speaker 2, etc.).” | Speaker Diarization: Analyzing multi-person interactions, like identifying who spoke how often in a meeting. |
| Video | “Transcribe the narration, noting any non-speech events like ‘[laughter]’ or ‘[applause]’.” | Accessibility and Context: Providing a complete text representation of the audible experience for those who cannot hear it. |
¶ Key Features
- Timestamp Precision - Get exact start and end times for each word or phrase
- Speaker Identification - Distinguish between multiple speakers in conversations
- Non-Speech Annotations - Capture ambient sounds like laughter, applause, or music
- Custom Formatting - Apply specific punctuation, capitalization, or formatting rules
¶ When to Use Other Work Types
- Need to identify temporal events? → Event Tagging
- Need to extract entities from text? → Named Entity Recognition
- Need to categorize audio content? → Classification
¶ Supported Data Types
- Audio - MP3, WAV, FLAC, and other audio formats
¶ Getting Started
- Define your transcription requirements (timestamps, speakers, formatting)
- Specify any domain-specific terminology or formatting rules
- Prepare your audio files
- Submit batches via API
- Track progress and download text transcripts
¶ How to Submit Batches
Indirect Demand (CSV File) - Best for large-scale projects with thousands of audio files.
Direct Demand (Inline Data) - Best for dynamic tasks or real-time processing.
¶ Best Practices
- Audio Quality - Ensure clear audio with minimal background noise for best results
- Speaker Context - Provide context about number of speakers and roles when possible
- Domain Terminology - Share industry-specific terms or jargon for accurate transcription
- Formatting Guidelines - Specify punctuation, capitalization, and timestamp requirements upfront
- Quality Control - Review sample transcripts to ensure accuracy and consistency