Named entity recognition: How NER works in practice

Speech waveform resolving into person, location, date, and money entities
Speech waveform resolving into person, location, date, and money entities

Named entity recognition (NER) turns spans of unstructured text into typed data, such as a person's name, an organization, a location, or a date. It is a core natural language processing (NLP) task, but a production pipeline needs more than a model prediction: clear labels, boundary rules, validation, and evaluation against real inputs. Here is how NER works, how the main approaches differ, and how to use its output safely in text and voice systems.

What is named entity recognition?

An NER result has two parts: the exact text span and its assigned type. Given “Jordan Lee works at Acme Health,” a system might mark “Jordan Lee” as PERSON and “Acme Health” as ORG. The boundary matters as much as the label because downstream systems need to know which source text supports the prediction.

The output is structured data that another system can search, count, validate, link to a record, or use in a workflow. Common entity types include people, organizations, locations, dates, times, products, and monetary values. A domain-specific system might instead recognize medications, named parties, courts, statutes, account types, airport codes, or appointment providers.

The schema matters. “October 12” can be a DATE in one project and part of an APPOINTMENT_TIME entity in another. NER does not discover a universal ontology. It applies the labels defined by a dataset, model, or rule set. Statistical NER systems also make predictions rather than guarantees, and their behavior depends strongly on their training examples. The spaCy NER guide illustrates both typed spans and this dependence on training data.

A worked NER example

Consider this input:

Jordan Lee will meet Acme Health in Phoenix on October 12.

A recognizer could produce:

Text spanTypeCharacter offsets
Jordan LeePERSON0–10
Acme HealthORG21–32
PhoenixGPE36–43
October 12DATE47–57

The offsets refer back to the original string using a start-inclusive, end-exclusive convention. Keeping offsets makes the result auditable and lets an interface highlight the exact evidence.

Many training datasets express the same spans as token labels. In the BIO scheme, B marks the first token of an entity, I marks the following tokens, and O marks tokens outside entities. “Jordan Lee” therefore becomes B-PERSON, I-PERSON. BIOES adds explicit end and single-token tags; Stanza's NER documentation shows the encoding token by token.

How named entity recognition works

One useful production extraction workflow has five stages. The exact architecture can differ, and span detection is the core NER step. The surrounding stages prepare the input and turn predictions into usable application data.

  1. Define the entity schema. Decide which types matter, what counts as a span, and how to annotate ambiguous, nested, or partial mentions. Write these decisions down before labeling data.
  2. Prepare the input. Split text into tokens or subword units while preserving alignment with the source. For speech, automatic speech recognition (ASR) produces the transcript first.
  3. Detect and type spans. Rules or a learned model assign labels to tokens or propose complete spans. A decoder then joins compatible token labels into entities.
  4. Resolve, normalize, and validate. Identify missing context before converting surface text into an operational value. “Friday at two thirty” still needs an AM/PM value, a date reference, and the caller's timezone before it can become a timestamp. “Market Street clinic” might resolve to an internal clinic ID.
  5. Use the result. Pass validated values to search, analytics, a database query, a knowledge graph, or an application workflow.

Stages four and five sit outside classic NER. This workflow is an application design, rather than a universal NER architecture. A model can correctly recognize Friday at two thirty as a time expression while the application still lacks the date, AM/PM value, timezone, or availability required to schedule an appointment.

The main NER approaches

Rule-based, traditional statistical, and transformer-based systems offer different tradeoffs. Transformers are also statistical machine learning models. We separate them here because their contextual representations, data needs, and serving profile differ from earlier sequence models.

Rule-based recognition

Rule-based NER uses exact dictionaries, regular expressions, token patterns, or grammar cues. A financial workflow might recognize an account number with a strict pattern and match product names against a controlled catalog.

Rules are easy to inspect and can deliver high precision for stable, closed sets. They become expensive to maintain when spelling, phrasing, or context varies. A dictionary also cannot decide reliably whether “Jordan” is a person or a place without contextual logic.

Rules remain useful beside a model. For example, spaCy's EntityRuler component can run alone or combine token and phrase rules with a statistical recognizer.

Statistical sequence models

Traditional statistical systems learn from labeled sequences. Depending on the model and configuration, hidden Markov models, maximum-entropy models, and conditional random fields (CRFs) may use features such as the current token, neighboring words, capitalization, prefixes, part-of-speech tags, or dictionary membership.

A CRF scores a sequence of labels jointly. That lets it learn transitions such as I-PERSON usually following B-PERSON, instead of classifying every token independently. These models can be compact and fast, and their engineered features can work well in a narrow, stable domain. Their main cost is feature design and weaker transfer when language or domain changes.

Transformer-based NER

A transformer encodes each token in the context of the surrounding text, then a token-classification or span-classification head predicts entity labels. Pretrained encoders can be fine-tuned on a project's labeled examples. The Transformers token-classification guide shows this standard setup and the alignment needed when one word becomes multiple subword tokens.

Transformers can use contextual signals that a name list lacks. They still inherit the project's schema, training data gaps, and annotation errors. They also add model size, serving cost, and latency considerations.

Hybrid and generative extraction

Many practical systems combine approaches: a learned model proposes spans, rules protect known identifiers, and validators reject impossible values. A hybrid design can also use a catalog to normalize model output without forcing the recognizer to memorize every current catalog item.

Large language models can return structured fields from prompts or examples. That can help when the schema changes often or labeled data is scarce. It is a different operating contract from classic span NER. A generative model may normalize, infer, or paraphrase a value instead of returning exact source offsets. Use schema validation and preserve source evidence when an extraction can trigger an action.

NER compared with adjacent language tasks

These tasks often share a pipeline, but they answer different questions.

TaskQuestion answeredExample output
Named entity recognitionWhich exact spans belong to which types?Phoenix → GPE
Intent classificationWhat is the utterance trying to accomplish?“Move my appointment” → RESCHEDULE_APPOINTMENT
Entity linkingWhich unique record does a mention refer to?Phoenix → knowledge-base ID for Phoenix, Arizona
Speech recognitionWhich words were spoken?Audio → “meet in Phoenix”
Relation extractionHow are two recognized entities related?Jordan Lee → works_for → Acme Health
Slot fillingWhich application fields are known or still missing?new_time=14:30, appointment_id=missing

Intent classification usually assigns one or more labels to the whole utterance. NER operates at the span level. One utterance can express a single intent and contain several entities.

Entity linking begins after a mention has been detected. It resolves that mention to a unique database or knowledge-base identifier. NER can label “Washington” as a location; linking decides whether it means the state, the city, or another record.

ASR converts audio into a transcript. NER then works on that transcript, unless a system uses a joint speech-and-entity model. An ASR error can remove or change the entity before NER sees it, so the two stages need separate and end-to-end evaluation.

Common NER use cases

NER is useful when text contains repeated facts that downstream software needs in a stable shape.

  • Search and content enrichment: index documents by people, organizations, locations, products, or events, and use those fields for filters and retrieval.
  • Document processing: extract named parties, courts, statutes, dates, amounts, or identifiers from contracts, invoices, filings, and support records.
  • Knowledge graphs and monitoring: detect mentions first, then link them and extract relations to build structured records.
  • Domain research: recognize entities such as drugs, genes, conditions, legal concepts, or financial instruments using a domain-specific schema.
  • Privacy workflows: locate candidate personal identifiers for review or redaction. NER alone is insufficient for a complete privacy policy because the required identifiers and acceptable error rates depend on the system and jurisdiction.
  • Conversational and voice systems: extract names, places, dates, account types, order numbers, or appointment details from user turns.

How to evaluate an NER system

Evaluate spans against human-labeled ground truth from the target domain. Under strict entity-level scoring, a prediction is a true positive only when both its boundary and type match the reference. “Acme” predicted as ORG does not exactly match the labeled span “Acme Health.”

The standard metrics are:

  • Precision: correct predicted entities divided by all predicted entities.
  • Recall: correct predicted entities divided by all reference entities.
  • F1: the harmonic mean of precision and recall.

The CoNLL-2003 shared task helped establish exact entity-level precision, recall, and F1 as common NER reporting measures. One overall F1 score is still too coarse for a production decision. Report at least:

  • results by entity type, with the number of examples for each type;
  • boundary errors separately from wrong-type errors;
  • performance on unseen names, abbreviations, misspellings, long entities, and ambiguous mentions;
  • domain and channel slices, such as email, chat, and ASR transcripts;
  • downstream success, such as the percentage of appointment requests completed with the correct validated record;
  • latency, compute cost, and abstention or confirmation rates at the intended traffic profile.

Use a held-out test set that reflects deployment traffic. Deduplicate near-identical examples across training and test data. Re-run the evaluation when the schema, domain vocabulary, upstream tokenizer or ASR system, or model changes.

A voice-agent example from speech to action

Voice systems expose a key production lesson: recognizing a span is only one step between audio and a safe action.

Assume the conversation date is October 5, 2026, and the caller's verified timezone is America/Los_Angeles.

  1. Spoken utterance: “Move my appointment with Dr. Maya Chen at the Market Street clinic from Thursday at nine to Friday at two thirty.”
  2. ASR transcript: move my appointment with doctor maya chen at the market street clinic from thursday at nine to friday at two thirty
  3. Intent result: RESCHEDULE_APPOINTMENT, produced by an intent classifier rather than NER.
  4. NER output under this example's schema:
    • doctor maya chen → PROVIDER
    • market street clinic → CLINIC
    • thursday at nine → DATE_TIME; application role: source appointment time
    • friday at two thirty → DATE_TIME; application role: requested appointment time
  5. Ambiguity resolution: the time expressions omit AM/PM, and the weekday names need to be grounded to the intended calendar week. Ask, “Do you mean Thursday, October 8 at 9 a.m. and Friday, October 9 at 2:30 p.m.?” Continue only after the caller confirms both dates and times.
  6. Normalization: after confirmation, resolve the expressions to 2026-10-08T09:00:00-07:00 and 2026-10-09T14:30:00-07:00 using the known conversation date and timezone.
  7. Validation: look up the authenticated caller's appointments. Confirm that the provider, clinic, and normalized source time identify one record, then check that the target time is available. Ask another focused question if the transcript or lookup remains ambiguous.
  8. Downstream action: summarize the proposed change, obtain final confirmation, submit the reschedule request, and record the original transcript, extracted spans, resolved IDs, and result for traceability.

DATE_TIME describes what each recognized span contains. The words from and to, the rescheduling intent, and application logic establish the source and requested-time roles. Those roles are part of relation extraction or slot filling, rather than universal NER types.

Spoken input creates errors that clean written-text tests miss. ASR can alter names, omit punctuation and capitalization, or split a phrase differently. Research on NER in spoken-dialog systems documents these missing written-language cues, while an ACL study of ASR and NER errors shows why conversational transcripts need more than a single aggregate F1 score.

Failure modes and practical fixes

Failure modeExamplePractical response
Ambiguous mentionJordan can be a person or placeAdd surrounding context and evaluate ambiguity as its own slice
Wrong span boundaryAcme instead of Acme HealthTighten annotation rules and measure boundary errors
Domain shiftA news model sees clinical abbreviationsFine-tune or retrain on representative domain text
Surface variationN.Y.C., NYC, New York CityAdd varied examples, then normalize after recognition
Nested or overlapping entitiesA location appears inside an organization nameChoose a model and output schema that support the required overlap
Inconsistent labelsAnnotators disagree on PRODUCT versus ORGMaintain written guidelines and adjudicate disagreements
ASR corruptionA provider name is mistranscribedUse speech-domain data, ASR alternatives when available, and confirmation for consequential fields
Stale catalogA new clinic is absent from a lookup listSeparate recognition from a versioned, regularly updated resolver

Confidence scores can support thresholds, but a raw model score is not automatically a calibrated probability. Set confirmation and abstention policies from measured errors and the cost of a wrong action. A miss in document analytics has a different impact from changing the wrong appointment or sending money to the wrong account.

Choosing an implementation

Start with the output contract, then choose the least complex approach that meets it.

  • Use rules or dictionaries for closed vocabularies, stable identifiers, and patterns with a clear grammar.
  • Use a pretrained NER pipeline for a fast baseline on common entity types. Test it on your own data before designing around its labels.
  • Fine-tune a statistical or transformer model when the domain, entity types, or phrasing differ materially from the pretrained data.
  • Use a hybrid pipeline when some fields are deterministic and others depend on context, or when recognized text must resolve against live business data.
  • Use generative extraction deliberately when flexible schemas outweigh the need for guaranteed source spans. Enforce structured output, retain evidence, and validate every operational value.

Before launch, define annotation rules, build a representative test set, set per-entity acceptance criteria, and decide what the application does with missing, conflicting, or low-confidence results. Log spans and downstream outcomes with appropriate privacy controls. Review failures by business impact, then update data, rules, models, or confirmation logic based on the error source.

NER is most useful when it has a narrow contract: find this evidence, assign these types, preserve the source span, and hand the result to validation.

At Dasha, we support two structured-data paths around a voice conversation. During a call, tools and functions use a team-defined JSON Schema for arguments sent to a webhook. After a call, post-call analysis applies configured labels to the transcript and returns structured results through a webhook or API. These are integration and analysis surfaces, separate from NER model training. Your technical team defines the schemas, business validation, downstream systems, and production acceptance criteria.

Start by classifying each required field as a during-call tool argument or an after-call analysis label, then configure the corresponding schema.

Related Posts

We use cookies for functional and analytical purposes. Please refer to our Privacy Policy for details.