An intent classification dataset can help you benchmark a model, bootstrap a prototype, or test a production routing policy. Those are different jobs. A public benchmark may prove that your pipeline works while missing the language, ambiguity, context, and failure costs of your own traffic. The right choice starts with the decision your system must make, then matches the domain, task shape, modality, languages, and out-of-scope behavior to that decision.
The short answer
Use a public intent classification dataset to compare approaches or validate an implementation. Use labeled examples from your own conversations to decide whether a system is ready for production.
Choose by these criteria, in order:
- Decision: What action will the predicted label control?
- Task: Is the output one intent, several intents, an intent plus slots, or an abstention?
- Domain: Do the labels and utterances resemble your workload?
- Input: Will the model see typed text, automatic speech recognition transcripts, audio, or conversation history?
- Coverage: Does the data include unsupported requests and difficult neighboring intents?
- Rights: Do the license, consent, privacy, and retention terms fit the intended use?
A large dataset with the wrong label system is usually less useful than a smaller dataset whose boundaries match the product decision.
What an intent classification dataset contains
Intent classification maps an utterance to a label representing the user's goal. For example, “My card still hasn't arrived” may map to card_arrival, while “Someone used my card” may map to card_payment_not_recognised. The label should correspond to a meaningful next step such as retrieving information, calling a tool, entering a workflow, or escalating.
An intent dataset usually contains:
- the original utterance or transcript;
- one or more intent labels;
- optional entity or slot spans, such as a date, location, or account type;
- a defined train, validation, and test split;
- optional conversation, speaker, language, or domain metadata; and
- optional out-of-scope examples that belong to no supported intent.
The term covers several different machine-learning tasks.
| Task shape | Output for one utterance | Use it when |
|---|---|---|
| Single-label classification | Exactly one intent | One route must own the request |
| Multi-label classification | Two or more nonexclusive intents | One turn can request several actions |
| Joint intent and slot filling | Intent plus structured fields | The route and its arguments must be extracted together |
| Hierarchical classification | Domain, then intent or sub-intent | A large taxonomy has meaningful parent-child structure |
| Out-of-scope detection | Supported intent or abstain | The system must reject, clarify, or escalate unsupported requests |
Out-of-scope (OOS) and out-of-distribution (OOD) are often used interchangeably in dataset descriptions, but they expose different risks. An OOS request has no valid label in your taxonomy. An OOD input differs from the data distribution, perhaps because of a new language, channel, customer group, or recognition error. An utterance can be in scope and still be OOD.
Six public intent classification datasets
No dataset below is a universal training set. Each isolates a different evaluation problem.
| Dataset | Domain and scope | Languages and input | Task type | OOS or OOD data | Access and license |
|---|---|---|---|---|---|
| CLINC150 | Task-oriented assistant, 150 intents across 10 domains | English text | Single-label intent classification | Dedicated out-of-scope train, validation, and test examples | UCI Machine Learning Repository download, Creative Commons Attribution 4.0 (CC BY 4.0) |
| BANKING77 | Online banking support, 77 fine-grained intents | English text | Single-label intent classification | No dedicated OOS class | PolyAI dataset card, CC BY 4.0 |
| MASSIVE 1.1 | Virtual-assistant requests, 60 intents and 55 slot types across 18 domains | 52 language varieties, text | Intent classification plus slot annotation | No dedicated OOS class | Amazon Science dataset card, CC BY 4.0 |
| MInDS-14 | E-banking requests, 14 intents | 14 language varieties, audio and transcripts | Spoken intent classification | No dedicated OOS class | PolyAI dataset card, CC BY 4.0 |
| NLU Evaluation Data | Human-assistant interactions across 21 home and personal-assistant domains | English text | Intent and entity annotations | No dedicated OOS class | Author repository, CC BY 4.0 |
| MIntRec2.0 | Multi-party dialogue from three television series, 30 in-scope intents | English text, audio, and video with dialogue context | Multimodal intent recognition | 5,700 out-of-scope utterances alongside 9,300 in-scope utterances | Zenodo v3 feature/text bundle: CC BY-NC-SA 4.0; author GitHub/Google Drive raw-video distribution: no file-level terms; THU-IAR Hugging Face mirror: CC BY-SA 4.0 |
The dataset license and the surrounding code license can differ. The terms attached to the data govern data reuse.
Pick CLINC150 for abstention experiments
CLINC150 is a good starting point when unsupported requests matter. Its full split contains 150 in-scope classes plus deliberately separate out-of-scope examples. It also offers smaller and imbalanced variants. That makes it useful for testing the tradeoff between accepting supported requests and abstaining on unsupported ones.
It does not prove that a threshold will work on your traffic. Its OOS examples reflect the benchmark's collection process, while production requests can sit much closer to a supported class.
Pick BANKING77 for fine-grained neighboring intents
BANKING77 puts 77 labels inside one banking domain. Labels such as card-payment problems, transfer problems, and cash-withdrawal problems create realistic confusion sets. It is useful for testing whether a classifier can separate semantically close routes.
The narrow domain is its strength and its limit. A strong result says little about a healthcare, travel, retail, or general-assistant taxonomy. The dataset also lacks a designated OOS split.
Pick MASSIVE for multilingual intent and slot work
MASSIVE 1.1 pairs intent labels with slot annotations across 52 language varieties. Its parallel structure supports cross-lingual comparison, while its 18 assistant domains cover a broader surface than a single-domain corpus.
Localization is not the same as naturally occurring production traffic. Use it to compare multilingual methods, then evaluate on conversations collected from each deployed locale and channel.
Pick MInDS-14 for spoken-input experiments
MInDS-14 contains audio, transcriptions, English translations, language identifiers, and intent labels for e-banking requests. It helps separate errors introduced by speech recognition from errors in the intent classifier.
Its 14-intent banking taxonomy is compact. It is a useful speech benchmark, not a substitute for audio and transcript conditions from your own phone routes, microphones, accents, and noise profiles.
Use NLU Evaluation Data for joint intent and entity pipelines
This corpus contains 25,716 annotated utterances from human-assistant interactions. The original release includes intent and entity-type annotations and the authors' annotation guidelines. That makes it useful for testing a pipeline in which the route and structured values are both part of the expected result.
Its home-assistant scenarios and collection design may be far from a commercial support or transactional workflow. Keep that domain gap visible in any interpretation of the score.
Use MIntRec2.0 only when multimodal dialogue is the problem
MIntRec2.0 adds dialogue order, speaker identity, audio, and video. It is useful for research on context and nonverbal signals, including OOS detection in multi-party conversations.
The source material is scripted television dialogue. License and access terms differ by distribution: the Zenodo v3 feature/text bundle includes a CC BY-NC-SA 4.0 license file; the author's raw-video download publishes no file-level license terms; and the THU-IAR Hugging Face mirror is labeled CC BY-SA 4.0. These labels should not be applied across distributions. MIntRec2.0 is a poor default for task-oriented routing, but a useful counterexample to text-only, isolated-utterance benchmarks.
How to choose a dataset for your use case
1. Define the label's operational consequence
Start with the action, not the corpus. Write down what changes when the model predicts each label. If cancel_transfer and check_transfer_status invoke different permissions or tools, they need a decision boundary the data can support. If two labels always trigger the same behavior, splitting them may add confusion without product value.
High-risk actions may require an abstain or confirm path even when the model has a top prediction. Intent confidence is evidence for a routing decision, not permission to perform an irreversible action.
2. Match the task shape
Do not force multi-intent requests into a single-label benchmark if the product must handle both actions. “Freeze my card and tell me which payment was declined” needs a multi-label or decomposed workflow. Likewise, a date, amount, destination, or identity attribute belongs in a slot or tool-argument evaluation rather than being hidden inside a coarse intent label.
If an agent uses a large language model (LLM) instead of a dedicated classifier, the data contract still applies. Version the label definitions, prompt, model, structured-output schema, and abstention rule together. The evaluation target is the complete decision system.
3. Match the deployed input
Text typed into a support widget is different from an automatic speech recognition transcript. Spoken input contains repairs, fragments, filler, overlap, recognition substitutions, and missing punctuation. A voice system may also need prior turns to resolve “yes, the second one.”
For voice, keep the original audio, reference transcript, recognition transcript, conversation context, and final intent linked where consent and retention policies allow. That structure shows whether a failure began in audio capture, transcription, context handling, or classification.
4. Treat a public corpus as a benchmark, not product coverage
A public dataset can help you reproduce a baseline, compare models, test a data loader, or measure few-shot behavior. It rarely represents your exact taxonomy and traffic.
Build the production evaluation set from approved, de-identified examples of the real workload. Include common requests, rare high-impact actions, adjacent intents, unsupported requests, short answers, corrections, language variants, and recognition failures. Synthetic examples can fill a planned scenario grid, but they should not replace a holdout of naturally occurring conversations.
Build a trustworthy production dataset
Write labels as decision rules
Every label needs a short definition, positive examples, exclusions, and a rule for ambiguous cases. Annotators should know whether to choose one label, several labels, request clarification, or mark OOS.
Audit confusion pairs explicitly. If annotators repeatedly disagree between cash_withdrawal_not_recognised and cash_withdrawal_charge, the problem may be missing context or unclear guidelines rather than model capacity.
Measure balance without flattening reality
Class counts matter, but equalizing every class can distort the production mix. Maintain two views:
- a natural-frequency set for expected overall performance; and
- a balanced or risk-weighted slice set that keeps rare intents visible.
Report both. A traffic-weighted average can hide a failing rare class, while an artificially balanced score can overstate the effect on typical traffic.
Find label noise before tuning the model
Review duplicates with conflicting labels, examples that contain too little context, templated paraphrases, stale intents, and annotations that violate the written rules. Sample both model errors and apparently correct high-confidence predictions. The second group can reveal systematic label mistakes that an error-only review misses.
Track disagreement rather than resolving it silently. Persistent disagreement is a signal to merge labels, add context, introduce multi-label output, or define an abstention path.
Prevent train-test leakage
Random row splitting is unsafe when near-duplicates come from the same template, conversation, customer, or generated paraphrase family. Group related examples into one split. For a changing production workload, hold out a later time window as an additional test.
Keep the final test set unavailable to prompt authors, model tuning, threshold selection, and retrieval. Once examples influence a change, they belong in regression data and a fresh untouched holdout should carry the release decision.
Protect conversation data
Production transcripts can contain names, phone numbers, account details, health information, payment data, and secrets spoken to an agent. Collect only fields needed for the task. Apply the approved consent, redaction, access, retention, and deletion controls before examples enter labeling or model pipelines.
Public availability does not make a dataset suitable for every use. License terms, source consent, geographic restrictions, and sector rules remain part of the data design.
Evaluate the system, not just the classifier
Use metrics that match the output:
- Single-label: accuracy, macro F1, per-intent precision and recall, and the confusion matrix.
- Multi-label: per-label precision and recall, micro and macro F1, and exact-match rate.
- Intent plus slots: intent accuracy, slot span F1, and complete-frame accuracy.
- OOS: OOS precision and recall at the deployed threshold, plus in-scope coverage and accuracy at that threshold.
Always segment results by language, channel, intent, customer or product group, input length, and other conditions that can change expected behavior. Attach cost to mistakes. Routing a balance question to a general FAQ is different from treating a cancellation as a balance question.
Offline classification metrics are only one layer of a conversational system. The release decision should also cover tool correctness, final state, policy, speech quality, latency, recovery, and reliability. Our voice agent evaluation guide shows how to connect scenario data to scorecards and release gates.
After release, sample low-confidence predictions, OOS traffic, human escalations, corrected routes, and high-impact outcomes. Add confirmed failures to the regression suite. Preserve a fresh holdout so continual improvement does not become continual test-set contamination.
A note on Dasha's older intent workflow
The original 2021 version of this page described the Dasha Platform and Studio workflow of that period. It used custom intent data, managed model training, and DashaScript checks such as messageHasIntent. Those examples are historical product material. They should not be read as instructions for the current Dasha application.
The current Dasha docs focus on configuring agents through system prompts, language models, voices, tools, schedules, and webhooks. They do not document the same custom intent-classifier training workflow. The durable lesson from the older guide is data discipline: define meaningful labels, collect representative examples, test neighboring intents, and learn from real conversations.
Intent labels can still serve as routing outputs, analytics tags, or evaluation dimensions inside a production conversational AI product. When your dataset and release gate are ready, start building on Dasha with the current managed runtime and API path.



