AI in Clinical Trials: Evidence and a Safe Rollout Plan

Clinical research coordinator reviewing a bounded AI-supported trial workflow
Clinical research coordinator reviewing a bounded AI-supported trial workflow

AI can reduce manual work across protocol planning, recruitment, study conduct, and data review. In a clinical trial, speed has little value if it weakens participant protection or evidence reliability. The useful question is specific: which task should AI perform, how much influence should it have, and what proof does that exact use require? A lifecycle view makes those decisions concrete.

What AI in clinical trials means

AI in clinical trials is the use of machine learning, natural language processing, computer vision, or generative models to design, conduct, monitor, analyze, or report a study involving human participants.

That broad definition covers three distinct roles:

RoleExamplePrimary control question
AI runs part of the trial workflowDrafting documents, finding candidate records, scheduling visits, or prioritizing data queriesCan the computerized process be validated, supervised, and reconstructed?
AI generates evidenceDeriving an endpoint from sensor data, predicting a covariate, or contributing to a safety or efficacy analysisIs the model credible for this defined context of use?
AI is the intervention under studyTesting an AI diagnostic, decision-support system, or closed-loop treatment toolDoes the protocol evaluate the AI, its users, and the resulting clinical pathway?

Each role needs its own validation plan. A model that drafts an internal meeting summary has little influence on the evidence. A model that excludes a participant, recommends a dose, or derives a primary endpoint can affect participant safety and the trial result.

Lower-risk near-term uses keep AI inside a bounded task and leave an accountable person in control. Examples include assisted candidate review, administrative participant communication, document drafting with source checks, and risk-based data review. Independent eligibility, dosing, endpoint, and safety decisions carry a much higher evidence and governance burden.

At Dasha, our relevant role is the participant-communication layer. Technical teams can use our managed voice AI runtime for approved recruitment outreach, visit scheduling, reminders, routine logistics, and transfer to study staff. We do not determine clinical eligibility, interpret adverse events, or replace the investigator.

Where AI adds value across the trial lifecycle

A sound design starts with one workflow problem, then gives the model the least authority required to address it.

Protocol design, site selection, and feasibility

Language models can search prior protocols, compare eligibility criteria, identify inconsistent definitions, and draft structured text. Predictive models can estimate site performance or test candidate criteria against historical and real-world data.

These outputs are planning evidence. Historical performance reflects which investigators and populations had access to earlier trials, so a site-ranking model can repeat the same access pattern. Three related problems deserve separate review:

  • Population bias: The training and feasibility data may underrepresent the demographic and clinical characteristics of the population that will use the product.
  • Site-access bias: A model trained on incumbent sites, large academic centers, or one claims network can rank sites that already have data and infrastructure above community sites that reach different participants.
  • Missingness bias: A blank field may mean that a source did not capture a site, service, or patient group. Treating missing data as zero can turn limited data access into a low site score.

Site selection should optimize for feasible enrollment and a representative population. FDA's representative enrollment guidance includes demographic characteristics such as age, sex, race, ethnicity, and residence, plus clinical characteristics such as comorbidities, organ dysfunction, disability, and body-weight range. Those dimensions affect the generalizability of the result.

A 2024 fair-ranking study treated incomplete site-data modalities and diverse enrollment as a joint site-selection problem. We infer an important operating rule from that work: better data coverage is not evidence of better access to the intended population.

Before relying on a site or participant ranker, teams should document source coverage, examine why values are missing, and compare performance across sites and relevant participant groups. Review subgroup false negatives and false positives, including whether a low-data site or underrepresented group is disproportionately removed from human review. Include community and lower-data sites in feasibility review when they can reach the target population. Once enrollment starts, monitor referral, screening, screen-failure, enrollment, and retention rates by site and relevant subgroup. A strong aggregate enrollment rate can hide a representation failure.

Protocol authors still own the scientific rationale, estimand, endpoints, inclusion criteria, and operational assumptions. Model-assisted exploration can inform those choices when it is labeled exploratory. Prespecification applies when an AI-derived variable, pipeline, or analysis contributes to confirmatory evidence.

Patient identification and prescreening

Machine learning can search electronic health records, registries, and referral notes for possible matches, then show the relevant evidence to a coordinator. This is a useful form of human augmentation.

TrialGPT, an experimental large language model system developed by National Institutes of Health researchers, was evaluated on three public cohorts totaling 183 synthetic patients and more than 75,000 trial annotations. In a separate pilot, two physicians screened all 36 combinations of six semi-synthetic oncology cases and six trials. TrialGPT assistance reduced aggregate screening time across those patient-trial combinations by 42.6%. The pilot was small, and its authors kept medical experts in the loop. The TrialGPT study supports assisted screening rather than autonomous enrollment.

An AI match should usually create a review queue. Missing history, temporal details, ambiguous criteria, and incomplete exclusion data can all defeat a plausible-looking match. A qualified member of the study team makes the eligibility decision.

Recruitment, scheduling, and retention

Conversational AI can contact people from an authorized list, explain approved study information, collect prescreen answers, schedule a coordinator call, send visit reminders, and route questions to staff. Voice is useful for people who cannot or do not want to use another portal.

An institutional review board (IRB) is the committee that reviews research methods and materials to protect participants' rights and welfare. FDA guidance says IRBs should review recruitment methods and materials, including scripted first contact that collects sensitive information. The guidance discusses receptionist scripts and predates current voice agents. Applying it to an AI first contact is our inference: an agent that follows the script and collects the same information belongs in the IRB review package. The review should cover the script, data collected, recipients, storage, opt-out handling, and escalation path. The FDA recruitment guidance also says recruitment material should not imply that an investigational product is safe, effective, or certain to benefit the participant.

Data collection, quality monitoring, and safety review

AI can flag improbable values, detect inconsistent records, group similar queries, identify missing forms, and prioritize cases for monitor review. Models can also classify reports or surface possible safety signals.

Priority is different from disposition. Qualified staff assess the case. AI must not overwrite or obscure source records. Any automated correction or transformation must be authorized, validated for its intended use, attributed to the person or system responsible, justified, and reconstructable from retained data and metadata. Keep the source input, model and pipeline version, output, reviewer action, and subsequent correction.

Digital health technologies create another input stream through wearables, sensors, and software. FDA's remote data guidance treats the hardware and software as part of the data-acquisition system. AI downstream does not remove the need to establish that measurement, transfer, and analysis are fit for the trial's purpose.

Statistical analysis, endpoints, and synthetic controls

AI can derive features from images, waveforms, speech, or sensor data. It can support prognostic modeling, covariate adjustment, and external-control construction. These uses can carry high regulatory impact because their output may influence a safety or efficacy conclusion.

For a confirmatory analysis, the statistical analysis plan should define the input population, data pipeline, model version, acceptance criteria, missing-data handling, subgroup checks, and sensitivity analyses. Freeze and document the preprocessing pipeline and model before the analysis. A mid-trial change to a confirmatory model requires interaction with the applicable regulator and an amendment to the statistical analysis plan. Without that path, the affected work should be labeled exploratory or post hoc rather than confirmatory. The European Medicines Agency's AI reflection paper applies these expectations to high-impact late-stage analyses.

Exploratory model development can continue in a separate environment when it is clearly labeled and kept out of the confirmatory analysis. A digital twin or synthetic control remains a model-based source of evidence. It does not automatically substitute for randomization, concurrent controls, or regulator agreement.

Document and submission support

Generative AI can summarize meetings, extract structured fields, draft plain-language material, translate approved content, and prepare first drafts of routine documents. Its main failure mode is a fluent statement with no support in the source record.

Require source-linked output, a named reviewer, version history, and an explicit rule against silently filling gaps. Generated text stays out of final protocols, consent materials, safety narratives, and regulatory submissions until the accountable reviewer approves it.

Use risk to decide how much validation AI needs

FDA's draft guidance for AI used to support drug or biologic regulatory decisions defines model risk through two factors: the model's influence on the decision and the consequence of a wrong decision. Its seven-step framework moves from the question of interest and context of use through risk assessment, a credibility plan, execution, documentation, and a final adequacy decision. The FDA credibility framework is draft guidance, and its context-specific logic is useful for internal governance.

The three tiers below are an illustrative Dasha operating translation. They are not regulator-defined categories.

Illustrative Dasha tierTypical useRecommended baseline controls
Lower model influenceDrafting a nonbinding internal summary or scheduling from approved availabilityApproved sources, access control, sampled review, error logging, and rollback
Assisted decisionRanking candidate records or prioritizing monitoring queriesLocked test set, mandatory human review, subgroup analysis, false-negative and false-positive review, traceable records, and release approval
Participant or evidence affectingExcluding participants, recommending dosing, deriving an endpoint, or triggering a safety actionPrespecified context of use, independent validation, frozen confirmatory model, strict change control, failure pathway, clinical and statistical oversight, and early regulator engagement where applicable

Lower model influence does not mean lower privacy risk. A scheduling workflow can still expose sensitive health information. Data classification, consent, security, vendor agreements, retention, and cross-border transfer rules need their own assessment.

A production rollout plan

Treat the model, integrations, and human workflow as one system. A high-scoring model can still fail through a missing field, a bad handoff, an inaccessible interface, or an unrecorded update.

1. Write a narrow context of use

Name the user, input, output, decision, population, location, and downstream action. “Improve recruitment with AI” is too broad. “Rank already authorized referrals for coordinator review, without excluding any referral” is testable.

2. Map data and authority

Document every system that supplies or receives data. Minimize fields, separate identifiers where possible, and give the AI only the actions it needs. Record which person can override, approve, and stop the workflow.

3. Establish a baseline and acceptance criteria

Compare the AI-assisted workflow with the current process. Measure time per reviewed case, false negatives, false positives, completion, escalation, and staff corrections. For participant-facing systems, measure opt-outs, failed identity checks, transfer success, and accessibility failures. Break results down by relevant site, language, device, and participant group.

4. Validate the complete computerized system

The International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use (ICH) publishes Good Clinical Practice principles for trial conduct. In the United States, FDA issued ICH E6(R3) as final, nonbinding guidance for industry in September 2025. Under its recommendations, the responsible party should use risk-based validation that considers intended use, participant safety, data importance, reliability, interfaces, access, security, audit trails, backup, and change control. The final E6(R3) guidance also says a sponsor may transfer trial activities to a service provider, while ultimate responsibility for sponsor activities remains with the sponsor.

Validation should cover model failures and ordinary system failures: unavailable source data, a timeout, a duplicated event, a transfer that does not connect, a model update, and an incomplete record.

5. Run prospectively with limited authority

Start in shadow mode or with mandatory review. Use real workflow conditions and the intended user population. Promotion criteria should include operational outcomes and error severity, as aggregate accuracy alone is insufficient. A silent false negative in eligibility screening matters more than an easily corrected spelling error.

6. Monitor versions, drift, and human behavior

Keep a registry of prompts, models, retrieval sources, thresholds, integrations, and release dates. Monitor input drift, subgroup performance, overrides, escalations, failed transfers, and unexpected uses. Retraining, a model-provider update, or a prompt change can alter the validated system and should enter change control.

A safe voice AI pattern for participant communication

We recommend that customers configure a trial voice agent outside eligibility and clinical decision systems. It should use sponsor- and IRB-approved scripts and contact lists only. Eligibility, dosing, adverse-event assessment, and every clinical decision stay with qualified study staff.

Give the agent a small set of approved tools, such as get_available_slots, book_coordinator_call, record_opt_out, and transfer_to_site. Customer-operated webhook services should authenticate and authorize each request, enforce the allowed state changes, and persist audit metadata for the call, tool request, model or workflow version, timestamp, result, and handoff. The model manages the conversation. Deterministic services own permissions and state changes.

Investigator-owned clinical decisions are separated from bounded AI communication and staff transfer

With Dasha, teams can expose bounded business actions through webhook-based tools and configure a warm or cold transfer to study staff. A recommended customer configuration should:

  1. identify the study and disclose that the caller is an AI agent;
  2. confirm the approved contact and permission state before disclosing study details;
  3. stay within the sponsor- and IRB-approved script and contact list;
  4. record an answer as provided, without converting it into eligibility or another clinical judgment;
  5. capture possible adverse-event information according to the protocol and standard operating procedure, then trigger prompt safety-team notification and a staff transfer for assessment and any required reporting;
  6. transfer symptoms, consent questions, distress, uncertainty, and requests outside the approved script to qualified staff;
  7. use a warm transfer when staff need context or a cold/direct transfer when immediate routing is the approved path;
  8. use deterministic failover if a transfer fails, such as a priority staff alert, an open callback task, and approved safety instructions, so the interaction cannot end as a silent success; and
  9. write a structured, traceable outcome under customer-defined logging and retention controls, including call, tool, notification, and handoff identifiers.

In the United States, the Health Insurance Portability and Accountability Act (HIPAA) Privacy Rule protects most individually identifiable health information held or transmitted by covered entities and business associates. That information is protected health information (PHI). The HHS Privacy Rule summary explains the scope.

Teams must confirm the current agreements and security posture with us before a Dasha workflow handles PHI. The customer configuration must also apply the approved data flow, access, logging, retention, residency, and subprocessor controls. If a required agreement or control is absent, keep PHI out of the workflow.

If the AI itself is being tested

When an AI system is the intervention, model accuracy is only one endpoint. The trial must capture the full human-AI pathway: who used it, what training they received, whether inputs were missing or poor quality, how outputs changed decisions, when users overrode it, and which errors reached participants.

Use the AI-specific reporting extensions alongside the core trial standard. SPIRIT-AI adds 15 items for protocols, while CONSORT-AI adds 14 items for trial reports. Both emphasize model version, input and output handling, user expertise, human-AI interaction, and error analysis. DECIDE-AI is useful earlier, when a system first enters a live clinical workflow and the team needs to evaluate safety, usability, and human factors before a large comparative trial.

What responsible progress looks like

AI adoption in drug development is real. FDA says its current approach draws on experience with more than 500 submissions containing AI components between 2016 and 2023. In 2026, FDA also announced two proof-of-concept trials that report defined endpoints and data signals to the agency in real time, and it solicited input on a broader pilot. That real-time trial initiative shows regulatory experimentation. It does not remove the protocol, monitoring, or validated data-path requirements.

The practical route is narrow: choose a measurable workflow, limit the model's authority, preserve source records, validate the whole system, and monitor every release. If participant communication is your bounded use case, talk to our team about configuring Dasha with narrow tools, explicit staff transfer, and customer-owned traceability controls.

Related Posts

We use cookies for functional and analytical purposes. Please refer to our Privacy Policy for details.