A negative genomic result should no longer automatically close the case.
While it closes the first analysis, it should also open a governed pathway for deciding whether, when, and how the case will be reviewed again.
The conventional genetic testing workflow is organized around a defined sequence of events: order the test, analyze the sequence, issue the report. An unresolved result often becomes a static record, even as gene-disease associations, variant classifications, phenotypic evidence, and analytical methods continue to evolve.
In June 2026, researchers at Boston Children’s Hospital, Harvard University, and OpenAI published an NEJM AI study on 376 previously unresolved rare disease cases. An AI-assisted reanalysis workflow surfaced leads that, after further review, testing, and clinical confirmation, contributed to 18 diagnoses. The additional diagnostic yield was 4.8%.
An AI-generated lead changes nothing for the patient until the surrounding workflow can identify the case, reconstruct the evidence, route it for specialist review, and return a confirmed result.
The study does not prove that AI can autonomously diagnose rare diseases. Rather, it shows that some unresolved cases contain evidence worth reconsidering, but the current service model does not consistently return to them. The 18 diagnoses are the result. The remaining 358 cases expose the operating challenge.
What Did AI Do, and What Remained a Clinical Decision?
The researchers applied OpenAI’s o3 Deep Research model to cases that had already undergone genomic testing and specialist review without a diagnosis.
For each case, the workflow used standardized Human Phenotype Ontology terms, occasional clinician notes, available clinical descriptions, patient metadata, and a filtered variant table. The model examined phenotype, inheritance, variant annotations, data quality patterns, and published evidence to propose possible molecular explanations.
The outputs were hypotheses for specialists to test.
At least two experts reviewed each candidate using the American College of Medical Genetics and Genomics and Association for Molecular Pathology framework. Disagreements were resolved through consensus. A finding counted as a diagnosis only after qualified specialists reviewed the evidence, classified the variant as pathogenic or likely pathogenic, obtained confirmation through a Clinical Laboratory Improvement Amendments-certified laboratory, and returned the result to the family.
The OpenAI summary of the study states the boundary directly: the model did not diagnose any patient or make any clinical decision.
The workflow widened the search and assembled evidence for review. Clinicians remained accountable for interpretation, confirmation, and communication.
That distinction should guide any diagnostics organization considering a similar capability. AI can help identify where specialists should look. It cannot assume responsibility for what enters the medical record or reaches a family.
What Does the 4.8% Yield Actually Tell Us?
The researchers reported the following results:

A yield below 5% can sound modest until the underlying population is considered. These were not new cases awaiting an initial analysis. Many had already been reviewed through commercial or institutional pipelines and discussed by multidisciplinary teams.
Even so, 4.8% should not be treated as a general benchmark for AI-assisted reanalysis.
Yield varied materially across cohorts. The early psychosis group produced the highest percentage, but it contained only 15 cases and therefore carried a wide confidence interval. Disease category, cohort selection, previous analytical methods, available phenotype data, and the likelihood of a monogenic explanation all affect the result.
Seven of the 18 diagnoses were rediscoveries. The diagnoses had already been established elsewhere but were absent from the local research record.
The rediscoveries expose a second failure mode. Some cases remain unresolved locally because clinically relevant information never returns to the active diagnostic workflow. Better reasoning alone will not solve fragmented records, unclear ownership of follow-ups, or disconnected reporting processes.
The study also did not measure time saved, cost, specialist effort, false-positive workload, or changes in patient care. A diagnostics leader cannot infer operating efficiency or return on investment from the 4.8% yield.
The result is a reason to run a proper evaluation. Costing that evaluation, and the program it might justify, is the next piece of work.
Why Are Diagnostics Programs Struggling to Revisit Unresolved Cases?
Reanalysis remains labor-intensive and often competes with incoming test volume for the same specialist capacity.
A 2024 study of Australian clinical and laboratory genetics services, published in the European Journal of Human Genetics, found that laboratories performed more than 25,000 new genomic tests between 2018 and 2021 but completed only 950 reanalyses during the same period. Genetics professionals identified workforce capacity as the principal barrier.
The study also found no consensus on cadence. Of the 121 respondents who answered the relevant question, 38% favored reanalysis every two to three years. Another 48% preferred an earlier interval, including 18% of all respondents who favored continuous reanalysis.
A meta-analysis cited in the paper reviewed 29 studies and reported an average additional diagnostic yield of approximately 10% after a median interval of about 24 months. The figure does not establish the expected return for every laboratory. It does show that unresolved cases can become diagnosable over time.
The reanalysis gap is a service-design problem (ownership, specialist capacity, consent, funding, patient recontact, and case prioritization). Sequencing is the part that already works.
AI may make a more responsive model feasible. Neither the Australian research nor the NEJM AI study proves that continuous AI monitoring works in clinical production. The former studied current practices and attitudes. The latter performed a one-time retrospective reanalysis.
Continuous reanalysis remains an operating model to test, for evidence has not established it as an outcome.
What Would a Governed Reanalysis Pathway Require?
A credible reanalysis program begins before the AI model is chosen. It needs clinical and operational rules that keep unsolved cases connected to new evidence. Seven matter most:
- Ownership. One team owns the unresolved cohort, including monitoring evidence, selecting cases, re-reporting, and recontacting. If everyone assumes someone else will reopen the case, no one does.
- Eligibility. Not every unsolved case belongs in the same queue. Explicit criteria decide which stay and when a case leaves the program.
- Reanalysis triggers. A calendar is only one trigger; a new gene-disease link, a reclassified variant, or fresh phenotype data should reopen a case too. AI can watch for these and specialists decide if they warrant review.
- Clinical governance. Every candidate runs a defined route: evidence, classification, confirmation, authorization, counseling, etc. Machine hypotheses and clinician findings stay separate. Model confidence is not evidence.
- Patient recontact. A new finding is worthless if it can't reach the patient. Consent, current contact details, and counseling capacity decide whether it does.
- Version control. Every reopened case must be reconstructable with build, sources, model version, evidence at the time, decisions taken. Otherwise, you know an answer changed but not why.
- Performance measurement. Yield alone is not enough. Track review time, false positives, confirmatory tests, and patients actually recontacted, all the measures show whether you built an improvement or another queue.
What Does Diagnostic Delay Cost?
The broader cost of delayed rare disease diagnosis is already substantial.
Research from the EveryLife Foundation for Rare Diseases found that the average diagnostic journey exceeded six years and involved close to 17 clinical encounters before a family received an answer. Across seven rare diseases studied, estimated avoidable costs associated with delayed diagnosis ranged from $86,000 to $517,000 per patient.
The study did not estimate how much AI-assisted genomic reanalysis would save. Many diagnostic delays arise before sequencing, and some unresolved cases remain beyond the reach of current medical knowledge.
Reanalysis addresses a narrower problem. It creates another opportunity for patients whose genomic data already exists but whose original analysis did not produce an answer.
The economic proposition must therefore be tested directly. A diagnostics organization needs to know whether prioritization reduces specialist workload, whether reopened cases produce actionable findings, and whether those findings change downstream care.
What Does Production Experience Teach Us About Genetic Testing Workflows?
In production diagnostics, the harder operational work begins after a plausible candidate has been identified.
Coditas encountered the same production principle while working with a US-based clinical diagnostics company running genetic testing programs across oncology, cardiology, neurology, and rare disease.
It involved rebuilding the operational system for one of the largest cancer programs in the country, running more than 500,000 genetic tests a year. Coditas connected their lab to more than 1,000 EHRs, cut case data assembly time from 45 minutes to 4 minutes (a 10x throughput gain with no added headcount), and shortened report turnaround time from 21 days to 11. The platform now handles more than 50,000 FHIR transactions a month at 99.97% uptime.
The numbers describe workflow orchestration, not genomic reanalysis, and should not be read as a predictor of reanalysis yield. The transferable lesson is the operating foundation. Production required interoperable clinical data, structured workflows, role-based access, traceability, and clear handoffs across the testing pathway — the same substrate any reanalysis program needs to carry an AI-generated lead from hypothesis to confirmed result.
A research model can surface a plausible candidate. A diagnostics program must carry that candidate through governed clinical operations without losing the evidence, accountability, or patient.
The Decision Diagnostics Leaders Need to Make
The NEJM AI study does not establish an autonomous diagnostic future. It exposes a weakness in the current testing model. A genomic report represents the evidence available when the case was analyzed. Treating an unresolved result as permanently closed allows the useful life of the analysis to expire before that of the genomic data.
Diagnostics leaders now have to decide whether an unresolved result remains a static report or enters a governed clinical pathway that can respond when relevant evidence changes. So, ownership. Who is accountable for the unresolved cohort, and what happens when an old case warrants another look? Naming that owner costs nothing, and it has to happen before any AI conversation even begins. Once that is defined, the operating model has to connect the data, review steps, permissions, and clinical handoffs to make that ownership executable.
At Coditas, we have already built the operating substrate behind the production results described earlier. That experience becomes relevant once a diagnostics program has made the ownership call and is ready to turn that decision into a working pathway.
I will be at HLTH USA in Las Vegas from November 15 to 18, Booth 4051, Level 2, Exhibit Hall, AI Zone. If your organization has not yet answered the ownership question, let’s make that the crux of our conversation.
FAQs
What Did the Boston Children’s AI Study Find?
Researchers used an AI-assisted workflow to reanalyze 376 previously unresolved rare disease cases. Following specialist review, further testing, and clinical confirmation, physicians established 18 diagnoses, producing an additional diagnostic yield of 4.8%.
Did AI Diagnose the Patients?
No. The model generated evidence-linked hypotheses. Qualified specialists reviewed the candidates, classified the variants, obtained laboratory confirmation, and made every clinical decision.
How Often Should Genomic Data Be Reanalyzed?
No universal cadence has been established. The appropriate interval depends on the disease area, patient population, consent, evidence changes, analytical capabilities, and available specialist capacity.
What Should a Genomic Reanalysis Program Measure?
The program should measure diagnostic yield, specialist review time, false-positive workload, time to candidate, confirmatory testing, revised reports, patient recontact, and changes in clinical management or care.

