AI in criminal justice can help people find records, transcribe speech, review large evidence sets, estimate risk, and draft routine material. Those uses carry very different consequences. A tool that retrieves candidate files is not equivalent to one that influences detention or sentencing. U.S. agencies need to define that boundary before procurement, validate each system in its real setting, preserve meaningful human authority, and give affected people a way to challenge consequential errors. This is educational information, not legal advice; requirements vary by jurisdiction and use.
What AI in criminal justice means
AI in criminal justice is an umbrella term for several kinds of systems used by law enforcement, prosecutors, defense teams, courts, pretrial services, corrections, probation, and parole. Treating them as one technology hides the differences that matter most.
- Machine-learning models classify records or estimate outcomes from historical patterns. Predictive-policing and some risk-assessment tools fall into this group.
- Computer vision and biometrics compare faces, fingerprints, images, video, or other sensor data. These systems often return candidates or similarity scores for further review.
- Risk-assessment models combine selected factors to estimate the likelihood of an event such as rearrest, failure to appear, or noncompliance. Some are statistical scoring systems rather than modern generative AI.
- Generative AI creates content. Proposed and emerging justice uses include summaries, draft reports, legal research, translation, and evidence queries.
- Speech systems transcribe or translate audio and can support voice interfaces. Their output requires review against the source.
The right risk category depends on the task, data, setting, and consequence. A transcript used to search a recording creates a different risk from a score considered at a detention hearing. The system around the model also matters: who may use it, which data it can reach, how its result is presented, and whether anyone can reject or correct that result.
Where AI is used across the justice lifecycle
The U.S. Department of Justice (DOJ) 2025 inventory lists 315 AI use cases across pre-deployment, pilot, deployed, and retired stages. That inventory shows breadth, but it is not evidence of 315 live deployments or 315 criminal-justice decision systems. Each entry needs its own status and scope check. DOJ's inventory makes those lifecycle distinctions explicit.
| Justice stage | Documented or emerging uses | What the evidence supports |
|---|---|---|
| Law enforcement | Face-search candidates, license-plate analysis, place-based forecasting, large-scale image or text search, and draft incident reports | The DOJ criminal-justice report documents existing biometric and predictive uses. Generative report drafting appears in a 2025 landscape study, where effectiveness evidence remains limited. |
| Forensic and digital evidence | Candidate retrieval from biometric databases, probabilistic DNA analysis, and assisted search across photos, video, files, and communications | The DOJ report describes selected real uses and limited adoption elsewhere. Qualified examiners still interpret results. |
| Prosecution and defense | Discovery search, legal research, document review, transcription, translation, and draft summaries or filings | The landscape study, sponsored by the National Institute of Justice (NIJ), identifies tools being developed, tested, piloted, or used. It does not prove that every listed workflow is deployed or effective. |
| Courts and pretrial services | Risk estimates that may inform release or supervision, plus emerging or potential research, summarization, and drafting assistance | The Bureau of Justice Assistance (BJA) documents pretrial risk-assessment use. The NIJ landscape study describes research, summarization, and drafting assistance as emerging or potential court workflows rather than established court-wide practice. Authority over release, supervision, and adjudication stays with accountable people under applicable law and policy. The federal judiciary's interim guidance specifically cautions federal courts against delegating core judicial functions to AI. |
| Corrections | Risk-and-needs assessment, classification, programming, case management, and proposed education or training support | The BJA overview documents risk tools for programming and management. Generative uses require separate evidence. |
| Probation and parole | Risk-and-needs assessment, supervision planning, and reentry support; hypothetical or emerging reminders, transcription, or translation assistance | Assessment and supervision planning are documented uses. Reminders, transcription, and translation are hypothetical or emerging examples here, not uses established by the BJA source. |
| Public administration | Non-emergency information, office routing, scheduling, and authenticated status retrieval | These are bounded workflow candidates. They do not establish authority for emergency response, legal advice, evidence evaluation, or a decision affecting liberty. |
This map separates a documented category from a proven outcome. Availability does not establish accuracy, fairness, time savings, or fitness for a particular jurisdiction.
AI can assist, but a qualified person decides
The safest operating model separates information assistance from legal authority.
| AI can assist with | A qualified person decides |
|---|---|
| Retrieve candidate records, images, or passages | Whether a candidate is relevant, reliable, or identifies a person |
| Transcribe, translate, or draft a summary | Whether the source is accurate and what evidentiary weight it deserves |
| Flag items for triage or suggest investigative leads | Whether to investigate, establish probable cause, arrest, or charge |
| Calculate a validated risk estimate | Detention, bail, sentence, supervision conditions, parole, or release |
| Surface authorities or draft routine material | Legal interpretation, credibility, admissibility, and adjudication |
| Route a non-emergency inquiry or retrieve an allowlisted status | Legal advice, emergency response, or any sensitive action on a person's case |
For face recognition, DOJ recommends treating search results as investigative leads. A result is not enough on its own to establish probable cause or positive identification without corroboration. The same report recommends that AI output never be the sole basis for a high-impact decision and that trained professionals review the output. DOJ's face-search guidance explains both limits.
A risk score is also probabilistic. It estimates outcomes for people with similar measured characteristics. It does not state with certainty what one person will do. The tool's calibration, meaning whether predicted rates correspond to observed rates, must be tested for the population and context where it will be used. The BJA validation guidance explains why context and multiple performance measures matter.
In September 2026, the federal judiciary reported that its interim guidance cautions courts against delegating core judicial functions, including decision-making and case adjudication, to AI. Judiciary users remain accountable for AI-assisted work. The guidance applies to the federal judiciary rather than every state system. Federal court guidance keeps the person responsible.
Benefits are conditional, not automatic
Well-scoped AI may speed up search across a large collection, produce a first-pass transcript, make routine information available in more formats, or standardize an administrative handoff.
Every benefit depends on a baseline and an acceptance test. Compare the assisted workflow with the process it replaces, measuring task completion, error severity, correction effort, downstream effects, and performance for relevant groups. Faster output has little value when a reviewer must reconstruct the source.
The distinction is especially important for generative AI. The 2025 NIJ-sponsored study found little empirical evidence supporting or refuting many promised product benefits and noted that early work had not established the expected efficiency gains for AI-assisted report writing. Its list of use cases is a landscape, not proof of effectiveness. The landscape study supports cautious pilots rather than broad outcome claims.
The main risks of AI in criminal justice
False matches and measurement error
Every model has an error profile. Face recognition can miss a true match or associate images of different people. The National Institute of Standards and Technology (NIST) reports that error rates vary with the algorithm, task, image quality, age, sex, and race. A benchmark result for one algorithm and one test condition cannot be generalized to every product or live search. NIST's demographic evaluation also shows why camera and image conditions belong in local testing.
The surrounding process can magnify error. A weak image, an overbroad gallery, or a reviewer who sees only the top result can turn a similarity score into false certainty.
Historical-data bias and population shift
Models learn from selected records. Arrest, incident, and supervision data reflect earlier reporting patterns, enforcement choices, data gaps, and human judgment. A model can reproduce those patterns even if it never receives a protected characteristic directly. The DOJ criminal-justice report explains that predictive-policing and risk-assessment data can contain gaps, errors, biases, or existing disparities that a model may carry forward.
Performance can change with the population, policy, collection method, or behavior. A tool validated elsewhere may not be calibrated locally. Local validation and periodic revalidation address different lifecycle points.
Opaque scores and weak challenge rights
A person may be unable to identify which data affected a score, whether the input was wrong, or how strongly the result influenced a decision. Proprietary restrictions can also prevent defense counsel, researchers, or an agency from examining the model and validation evidence.
When a risk tool contributes to detention, sentencing, supervision, or access to programs, opacity can raise due process or civil rights concerns depending on the facts, the system's role, and applicable law. The DOJ report notes that people subject to risk assessment may lack notice, information about the model and its performance, visibility into inputs, or an opportunity to correct mistakes. Notice without a usable explanation is thin protection. Affected people need an appropriate way to correct input data, add relevant context, and challenge consequential use.
Generative-model confabulation and source loss
Generative systems produce likely sequences, so a fluent answer can include an invented fact, quotation, citation, or connection. A summary can omit a qualification that changes the meaning of a record. Translation and transcription can also alter names, dates, negation, or speaker attribution. The NIJ landscape study documents fabricated legal cases, citations, and transcription content and calls for rigorous output validation.
Review must return to the authoritative source. If a system cannot preserve the link between each material statement and its source, it is a poor fit for evidentiary or legal work.
Privacy, confidentiality, and security
Justice data can include sealed records, criminal-history information, biometrics, health information, victim data, and attorney-client or work-product material. For generative AI, the NIJ landscape study warns that cloud-based tools may transmit, store, or retain inputs and that some free tools may be unsuitable for confidential or sensitive information.
An agency needs a data map covering collection authority, access, provider processing, retention, deletion, sharing, incidents, and use of submitted data. De-identification does not replace lawful authority and access control.
Automation bias and nominal human review
Reviewers can uncritically accept or over-trust AI-generated content, even in high-stakes settings. The NIJ landscape study identifies this as automation bias and recommends training and verification protocols. Adding a reviewer to a workflow does not solve that problem by itself. The reviewer needs training, time, source access, authority to reject the output, and an escalation route.
A safeguards and procurement checklist
The DOJ report provides nonbinding recommendations. The NIST framework overview describes the Artificial Intelligence Risk Management Framework (AI RMF) as voluntary and notes that version 1.0 is being revised in 2026. Neither replaces the Constitution, statutes, court rules, rules of evidence, collective-bargaining duties, records law, procurement rules, or local policy. The AI RMF Core organizes risk work around govern, map, measure, and manage, with continuous review across the lifecycle.
- Define authority, purpose, and prohibited uses. Name the legal and policy authority for the data and workflow. State who may use the system, for what task, and which outputs can never trigger action. Keep identification, credibility, probable cause, charging, detention, sentencing, parole, and adjudication outside an automated decision path.
- Complete impact and data assessments. Map affected people, foreseeable harms, non-AI alternatives, data provenance, retention, access, and downstream sharing. Include defense, civil-rights, privacy, security, operational, and community perspectives before procurement.
- Validate locally under deployment conditions. Test the exact version, data sources, thresholds, interface, hardware, and reviewer workflow. For risk tools, validate against the relevant local population and outcome. For speech or vision, include the audio and image conditions the system will actually encounter.
- Test performance across relevant groups. Report false positives, false negatives, calibration, failure-to-produce-a-result rates, and serious error types. Break out results for groups and operating conditions relevant to the use. One aggregate accuracy number is insufficient.
- Design meaningful human review and corroboration. Give reviewers training, time, source evidence, authority to reject, and a clear escalation route. Require independent corroboration before a face-search lead or generated assertion influences a consequential action.
- Provide notice and a usable challenge path. Where appropriate, tell affected people that a system was used, what role it played, which inputs mattered, and how to correct data, add context, seek review, or appeal. The DOJ recommendations call for notice, explanation, and an opportunity to respond in risk-assessment use.
- Keep versioned operational records. Log the model and policy version, authorized user, input source, output, confidence or uncertainty where meaningful, reviewer action, override, final decision, incident, and correction. Apply access and retention rules to the logs themselves.
- Contract for evidence and control. Require validation evidence for the purchased version, audit access, known limitations, data-use terms, security duties, incident notification, advance notice of material model or policy changes, and the ability to suspend an update. Avoid contracts that make independent evaluation impossible.
- Monitor outcomes and revalidate. Track drift, subgroup performance, overrides, complaints, incidents, and downstream outcomes. Revalidate after material changes to the model, population, policy, data, interface, or workflow. The BJA validation model illustrates why prediction quality must be checked against observed outcomes.
- Maintain a stop and rollback path. Set pause thresholds, preserve a safe manual process, assign incident authority, and rehearse rollback.
Where Dasha fits: bounded, non-emergency voice workflows
At Dasha, we provide technical teams with a platform for building and operating custom voice AI agents. Our current product supports configurable phone workflows, call transfers, webhook tools, call histories, testing, and completed-call inspection. These capabilities can support administrative access. They do not make Dasha a criminal-justice decision system.
The examples below are bounded workflow examples, not Dasha customer case studies:
- answer approved, non-emergency public-information questions from an allowlisted source;
- route a caller to the correct office or trained person;
- schedule an appointment or a human interview;
- send an administrative reminder;
- retrieve an allowlisted status from an authoritative system after appropriate authentication; or
- collect structured, disclosed intake before prompt human escalation of sensitive facts.
We do not position Dasha for 911 or emergency dispatch, suspect or witness interrogation, covert impersonation, credibility assessment, evidence analysis, legal advice, forensic conclusions, risk scoring, or decisions about probable cause, arrest, charging, detention, bail, sentencing, parole, prosecution, or adjudication.
The customer retains responsibility for workflow rules, system permissions, data handling, telephony, legal and compliance review, monitoring, and staffed escalation. Our voice AI infrastructure guide explains that ownership boundary. Dasha's webhook tools, call history, and Call Inspector provide implementation and review surfaces within that boundary.
A cautious pilot sequence
- Select one low-consequence, non-emergency administrative task with a clear manual fallback.
- Write the allowed topics, forbidden topics, authentication rules, disclosures, and immediate transfer conditions before building the agent.
- Connect only the minimum allowlisted sources and read-only tools needed for the task. Keep authoritative case data in the system of record.
- Test ambiguity, silence, noise, identity failure, sensitive disclosures, prompt manipulation, wrong data, tool timeouts, and transfer failure. Our voice-agent testing guide provides a failure-first structure.
- Run a limited pilot with trained staff available for escalation. Review calls, tool activity, caller outcomes, complaints, and near misses.
- Release only when the workflow meets its accuracy, privacy, escalation, and rollback gates. Stop when a severe incident or repeated boundary failure crosses the approved threshold.
Frequently asked questions
What are examples of AI in criminal justice?
Examples include face-search candidate generation, license-plate analysis, digital-evidence search, probabilistic DNA interpretation, risk-and-needs assessment, transcription, translation, legal research, and draft reports. Deployment and evidence vary widely.
Can AI make criminal-justice decisions?
AI can provide information or a probabilistic estimate, but it should not hold legal authority. Identification, probable cause, arrest, charging, detention, bail, sentencing, parole, credibility, evidentiary weight, and adjudication require accountable human decision-makers under applicable law and policy.
What are the biggest disadvantages of AI in criminal justice?
The main risks are false matches, unrepresentative data, population shift, opaque scores, invented generative output, confidentiality loss, automation bias, and weak mechanisms for challenging errors.
Is a criminal-justice risk score a prediction about one person?
It is an estimate based on outcomes observed for people with selected similar characteristics. It does not establish what one individual will do. Its use requires clear purpose, local validation, understandable limits, accurate input data, human judgment, and a way to contest material errors.
Will AI replace judges, lawyers, or justice personnel?
AI can reduce some search, transcription, drafting, and administrative work. It cannot assume professional accountability or lawful decision authority. The federal judiciary's current guidance keeps core judicial functions with courts and holds users accountable for AI-assisted work.
The right next step is a narrow pilot only after the agency has documented authority, prohibited uses, human escalation, validation criteria, and stop conditions. If those safeguards cannot be enforced, keep the workflow manual.



