Named Entity Recognition
Named Entity Recognition (NER) is the process of identifying and classifying specific proper nouns or standardized phrases within text. The output is the marked span of text and its corresponding label (e.g., Person, Location, Date).
NER is foundational for understanding the structured elements contained within unstructured text data.
Ideal for: Information retrieval, compliance analysis, brand monitoring, clinical text processing, knowledge base creation, and any scenario requiring structured data extraction from text.
¶ When to Use
Use Named Entity Recognition when your input is text and your goal is to identify and classify specific proper nouns or standardized phrases within that text.
¶ Simple Entity Identification Tasks
Ideal when you need to extract and categorize core entities from a document or communication.
| Input Type | Question Example | Purpose / Requirement |
|---|---|---|
| Document Text | “Mark the names of all people, organizations, and locations in this news article.” | Information Retrieval/Summarization: Indexing content to make it searchable by entity type. |
| Legal Text | “Identify all dates, monetary values, and contract identifiers in this legal agreement.” | Compliance and Contract Analysis: Extracting key structured data points from legal documentation. |
| Social Media Post | “Tag any mentioned product name or brand name in this customer review.” | Brand Monitoring/Sentiment Analysis: Tracking mentions of specific products or companies. |
¶ Contextual and Fine-Grained Entity Tasks
Use this for more complex scenarios where entities need to be classified into specialized categories or where the context is important for defining the entity.
| Input Type | Question Example | Purpose / Requirement |
|---|---|---|
| Medical Record | “Label all mentions of symptoms, drug names, and dosage amounts.” | Clinical Text Processing: Structured data extraction for research or electronic health records (EHRs). |
| Technical Manual | “Mark all instances of part numbers, error codes, and system components.” | Knowledge Base Creation: Building a structured index of technical terms and identifiers. |
| Dialogue Text | “Identify all entities, but also label the relationship between a Person and a Location (e.g., [Person] lives in [Location]).” | Relationship Extraction: A more advanced task built on top of NER to understand how entities interact. |
¶ Key Difference from Entity Extraction
Named Entity Recognition (NER) is often used as an alias for Entity Extraction, particularly when the task involves tagging spans of text with fixed labels (like Person or Location). The term Entity Extraction can be broader, sometimes including the extraction of associated attributes or normalization of the extracted data into a structured format (e.g., linking the text “J. Smith” to a specific ID in a database). For annotation purposes, they are frequently used interchangeably to mean span-based labeling.
¶ Common Entity Types
- Person - Names of individuals
- Organization - Companies, institutions, government bodies
- Location - Cities, countries, addresses, geographic features
- Date/Time - Temporal expressions
- Monetary Values - Currency amounts, financial figures
- Product Names - Brands, product identifiers
- Medical Terms - Drugs, symptoms, conditions
- Technical Terms - Part numbers, error codes, system components
¶ When to Use Other Work Types
- Need simple text categorization? → Classification
- Need document field extraction? → Document Annotation
- Need speech-to-text conversion? → Transcription
¶ Supported Data Types
- Text - Plain text, documents, articles, social media posts, chat transcripts
¶ Getting Started
- Define your entity types and classification categories
- Specify any domain-specific entity types (medical, legal, technical)
- Prepare your text data
- Submit batches via API
- Track progress and download structured entity annotations
¶ How to Submit Batches
Indirect Demand (CSV File) - Best for large-scale projects with thousands of text documents.
Direct Demand (Inline Data) - Best for dynamic tasks or real-time processing.
¶ Best Practices
- Clear Entity Definitions - Define precise criteria for each entity type
- Domain Terminology - Provide industry-specific entity examples and edge cases
- Ambiguity Guidelines - Document how to handle entities that fit multiple categories
- Boundary Rules - Specify how to mark entity spans (full names vs. partial)
- Quality Control - Review sample annotations to ensure consistency across entity types