Eye care produces structured images, repeated measurements, and time-sensitive referral decisions, which makes ophthalmology a strong field for narrowly defined AI. Strong study results can still fail when the camera, patient population, workflow, or handoff changes. A safe rollout starts by defining the task, its evidence and regulatory status, the permitted data flow, and the person who owns each exception.
What AI in ophthalmology means in practice
AI in ophthalmology covers several types of systems with different inputs, risks, and levels of autonomy.
| System type | Typical input | Typical output | Appropriate role |
|---|---|---|---|
| Imaging AI | Fundus photographs, optical coherence tomography (OCT), visual fields, corneal topography | Detection, segmentation, grading, measurement, or a referral recommendation | Screening or clinician support within a defined indication |
| Predictive AI | Images plus longitudinal clinical data | Estimated progression or treatment-response risk | Prioritization and care planning with clinician review |
| Generative AI | Notes, letters, policies, and patient questions | Draft summaries, nonclinical administrative logistics, referral letters, or documentation | Drafting and information retrieval from approved sources, with an appropriate reviewer. Patient-specific clinical instructions require qualified-clinician review plus intended-use and regulatory assessment. |
| Voice and workflow AI | A live conversation plus scheduling or practice-system access | Appointment action, reminder outcome, routed call, or staff handoff | Administrative workflows with limited permissions and defined escalation |
| Surgical research systems | Video, instruments, sensors, and device telemetry | Phase recognition, guidance, or physical assistance | Research and training. Clinical use depends on a specific device authorization. |
The first boundary is clinical versus operational. A retinal system can interpret an image and influence a referral. A voice agent can book that referral, explain logistics from approved content, and connect the patient with staff. The operational system should receive only the minimum status needed for its approved task. It should never infer a diagnosis from an image or turn a screening result into general eye-care advice.
At Dasha, we work on the operational side. Our voice agents can request tightly scoped business actions through webhook-based tools and route calls to staff through direct, warm, or HTTP transfers. This fits scheduling, reminders, intake routing, and staff handoff. We do not interpret eye images or replace an ophthalmologist.
Our public security page states that Dasha does not currently publish a HIPAA Business Associate Agreement.

Where clinical AI is authorized and where it remains research
Three labels keep clinical claims clear:
- Authorized means a regulator has permitted a specific device for a defined population, input, task, output, and use environment.
- Assistive means software supports a clinician inside its cleared or authorized indication. It does not gain autonomy because a clinician can review the result.
- Research means a study demonstrated performance under stated conditions. That evidence alone does not establish authorization or routine clinical readiness.
Autonomous diabetic-retinopathy screening has a narrow authorized use
The FDA authorized IDx-DR to detect more-than-mild diabetic retinopathy in a defined primary-care workflow. The authorized scope is adults age 22 or older who have diagnosed diabetes and no previous diabetic-retinopathy diagnosis, using retinal color images captured with the Topcon NW400, also identified as the TRC-NW400. It does not screen for diabetes, glaucoma, cataract, or other retinal disease.
The authorized pathway returns one of three results: more-than-mild diabetic retinopathy detected, more-than-mild diabetic retinopathy not detected, or insufficient quality. Insufficient quality is not a negative result. The labeling requires the patient to be immediately retested or referred to an eye-care professional when the system provides no result. If it still cannot generate a result after pharmacologic dilation, the patient should be seen by an eye-care professional.
The counts behind the performance metrics matter. The FDA-reviewed clinical study enrolled 900 participants at 10 primary-care sites, 892 completed all procedures, and 819 could be fully analyzed. Among the fully analyzable participants, observed sensitivity was 87.4% and observed specificity was 89.5%. The model-based primary sensitivity was 87.2%, and the enrichment-corrected primary specificity was 90.7%. FDA reported imageability as 96.1%, with a denominator of 819/852. The FDA De Novo summary documents the population, camera, outputs, fallback, and study accounting.
An operational system connected to this pathway should receive a limited state such as referral required, routine rescreening due, or insufficient-quality fallback. Its patient-facing language and next action must match the authorized device output. It should not receive the raw image or generate a broader clinical interpretation.
Most other ophthalmic examples remain research
Published studies show what may become useful, but each result remains tied to its study data, input, task, and reference standard.
| Area | Exact studied task and input | Maturity |
|---|---|---|
| Retinal OCT | Segment tissue and recommend referral urgency from OCT scans across retinal conditions in a two-stage system | Research. The OCT referral study did not authorize autonomous use across clinics or scanner types. |
| Glaucoma | Detect glaucomatous optic neuropathy from color fundus photographs | Research. This primary glaucoma study supports a defined image-classification task. |
| Cataract | Detect and grade lens opacity from slit-lamp and retro-illumination lens photographs | Research. The cataract model study does not establish a general cataract diagnostic service. |
| Cornea | Distinguish keratoconic from normal eyes using six color-coded maps from swept-source anterior-segment OCT | Research. The single-center accuracy study reported this exact task and hardware-derived input. |
| Pediatric ophthalmology | Classify plus disease from retinal photographs of infants undergoing retinopathy-of-prematurity screening | Research. The ROP study evaluated image classification, not autonomous management. |
| Neuro-ophthalmology | Detect papilledema from ocular fundus photographs | Research. The BONSAI study does not authorize generic neuro-ophthalmic triage. |
Retinal foundation models widen the research agenda. RETFound was pretrained on 1.6 million unlabeled fundus and OCT images and then adapted to ocular and systemic prediction tasks. The RETFound study supports a reusable research base. Each adapted clinical task still needs its own validation and applicable regulatory review.
Clinician review does not remove image software from device oversight
The FDA's January 2026 CDS guidance draws a firm boundary. Software that acquires, processes, or analyzes a medical image for a medical purpose remains a device function under the guidance. The non-device clinical decision support exclusion requires all four statutory criteria to be met, and medical-image analysis fails the first criterion.
Giving a clinician the original image, model output, or opportunity to review the recommendation can improve safe use. It does not convert image-analysis software into non-device software. Independent clinician review is a separate criterion for software that otherwise qualifies for the exclusion.
Why study accuracy can fall apart in a clinic
Sensitivity and specificity describe a model under stated conditions. A working service also depends on the camera, operator, patient, network, referral capacity, and follow-up process.
A deployment study observed diabetic-retinopathy screening across 11 clinics in Thailand. Researchers found tension between the model's image-quality threshold and the images staff could capture in a resource-constrained setting. Those failures changed nursing work and the patient experience despite strong underlying diagnostic performance. The clinic deployment study shows why an ungradable image is a workflow outcome, rather than a neutral technical event.
Five failure modes need explicit handling:
- Dataset shift. Performance can change with a new population, prevalence, camera, scan protocol, or referral setting.
- Poor acquisition. Small pupils, motion, media opacity, and operator technique can produce an ungradable image.
- Scope creep. Staff may treat a diabetic-retinopathy screening output as a general retinal exam.
- Automation bias. A plausible score or polished summary can suppress appropriate clinical skepticism.
- Broken follow-through. Correct triage has little value if the referral is not scheduled or the result never reaches the responsible clinician.
The safety unit is the full care pathway. Measure repeat imaging, referrals, completed visits, delayed cases, clinician overrides, and unresolved exceptions after the model produces its output.
Voice AI belongs on the operational side of the boundary
Ophthalmology practices field routine questions about appointments, preparation, transportation, referrals, and office policies. Voice AI can handle these structured administrative conversations while sending clinical questions and every symptom disclosure to qualified staff. The agent never decides whether a symptom is urgent and never de-escalates a symptom report.
A safe call flow separates routine service from urgent escalation:
- The agent states its role and uses the minimum identity data required for the approved task.
- Read-only tools retrieve approved locations, appointment types, nonclinical administrative logistics, and available slots from a source of truth.
- The patient confirms a choice before a separate write tool creates or changes an appointment.
- The agent confirms the booking only after the scheduling system returns success and a booking identifier.
- A non-symptom clinical question moves to the clinic's qualified-staff queue under its response-time policy. The agent does not answer it.
- Every symptom disclosure routes to qualified staff. The agent never decides whether it is urgent or de-escalates it. A static, clinic-approved emergency phrase list may only accelerate that escalation. When a phrase matches, the agent attempts an immediate transfer to qualified staff. If that transfer fails, it invokes the clinic-approved emergency fallback and raises the priority alert. It does not place the caller into a generic callback queue.
- The application records tool results, handoffs, opt-outs, failed transfers, and final disposition under the practice's access and retention rules.
Outbound artificial-voice calls also need a jurisdiction-specific compliance gate before dialing. That gate should determine the call's purpose, destination and number type, applicable consent or exemption, required caller identification, opt-out method and suppression action, and any call-recording notice or consent. If the configured rule set cannot establish a permitted path, the system should deny the call. There is no single consent or recording script that works for every jurisdiction and call purpose.
The FCC's declaratory ruling on AI voices confirms that AI-generated human voices fall within the Telephone Consumer Protection Act's artificial or prerecorded voice rules. It states that prior express consent is required absent an emergency purpose or exemption, that identification and disclosure rules apply, and that advertising or telemarketing messages must provide the specified opt-out methods.
Administrative AI does not remove healthcare privacy obligations. In the United States, a cloud provider that maintains or transmits electronic protected health information is generally a business associate even if it cannot view encrypted data, according to HHS cloud guidance. Encryption, access controls, retention, incident response, vendor agreements, and patient notices belong in the design before real patient data enters the workflow.
A seven-step rollout plan for an ophthalmology practice
1. Define one decision or action
Use a statement precise enough to audit. A clinical example is: “Run the authorized IDx-DR screening pathway for adults age 22 or older with diagnosed diabetes and no prior diabetic-retinopathy diagnosis, using Topcon NW400 images and the authorized result and image-quality fallbacks.” An operational example is: “Reschedule established-patient OCT follow-ups within approved appointment types.”
2. Name the owner and every failure path
Assign a clinical owner for diagnostic outputs and an operational owner for scheduling or communication. Define what happens after a positive, negative, insufficient-quality, disconnected, failed-tool, or failed-transfer result. Each path needs a responsible queue and response time.
3. Match evidence to the local pathway
For clinical software, record the authorized indication, population, imaging hardware, acquisition protocol, reference standard, exclusions, and fallback. Evaluate it using the local disease spectrum and image-quality distribution.
The current cross-regulator principles come from IMDRF's 2025 final Good Machine Learning Practice document, which FDA publishes on its GMLP page. They cover representative datasets, independent training and test sets, human-AI team performance, clear user information, and monitoring of deployed models across the total product life cycle.
4. Set the human boundary
Specify which outputs require clinician review, which low-risk administrative actions may execute automatically, and which static clinic-approved phrases accelerate symptom escalation. Generated instructions are limited to nonclinical administrative logistics. Patient-specific clinical instructions require qualified-clinician review plus an intended-use and regulatory assessment. Human review alone does not determine device status under FDA clinical-decision-support guidance. A qualified clinician owns interpretation of new symptoms and treatment decisions.
5. Integrate only the required sources of truth
Connect the system to imaging, health-record, scheduling, and referral systems only as required. Use constrained tool inputs, idempotent writes, role-based access, and durable event logs. A model should never infer that a booking, referral, or message succeeded after a timeout.
6. Run in shadow mode, then constrain the launch
First compare outputs with the current process without allowing the system to affect care. Review false negatives, false positives, insufficient-quality results, tool errors, and handoffs. Limit the live launch by site, device, appointment type, call purpose, schedule, or traffic share. Our voice-agent testing guide explains how scenario suites and trace review expose conversational failures before a wider rollout.
7. Monitor outcomes, drift, and workload
Clinical metrics should include sensitivity, specificity, insufficient-quality rate, referral completion, time to specialist review, overrides, and performance by relevant population and device. Operational metrics should include confirmed bookings, tool success, transfer success, abandoned calls, repeat contacts, opt-outs, staff corrections, and appointment show rate. Track the work created by exceptions alongside the work saved.
AI changes the workflow, while accountability stays with people
AI can support repeatable perception, measurement, retrieval, and routing. Ophthalmologists still combine incomplete evidence, examine atypical cases, discuss tradeoffs, perform procedures, and remain accountable for care. An authorization covers a defined system and workflow, while research findings show what may be possible next.
The safe design is narrow: clinical AI performs only its validated task, operational AI moves the patient through the next approved step, and a qualified person owns each exception. If you are building the operational voice layer, evaluate Dasha with a constrained scheduling or reminder workflow and synthetic data first.
