Soleil Research | Evidence Review
Large language models are increasingly proposed as clinical decision-support tools, but benchmark performance alone cannot show whether clinicians will use them effectively inside a real emergency department. A newly published Nature Medicine study prospectively evaluated SHAKED, a multi-LLM clinical decision-support system, during routine emergency care. The most important finding was not a dramatic improvement in throughput. It was a human-factors signal: clinician use fell substantially over only four weeks.
Plain-Language Summary
Researchers evaluated SHAKED in a tertiary emergency department over four weeks, analyzing 1,138 patients across two parallel clinical units. One unit had access to SHAKED while the other followed routine workflow.
Physician adoption of the system declined from 68% to 30% during the study. Use also fell as shifts progressed, consistent with workload-sensitive disengagement. Physicians were more likely to use SHAKED when seeking radiology consultation support.
Expert reviewers judged 99 of 100 sampled SHAKED outputs clinically appropriate, and the investigators detected no adverse events attributable to the system during the study. However, emergency-department length of stay was 4.9 hours in both wings, and the intention-to-treat analysis of consultation cycle time showed only a non-significant trend toward improvement.
The study therefore provides useful prospective evidence about feasibility and clinician behavior, but it does not establish that LLM decision support improves patient outcomes or is ready for routine clinical deployment.
Why This Study Matters
Medical AI is often evaluated by asking whether a model can answer clinical questions accurately. Real-world deployment raises a harder question: can the system fit into a complex clinical workflow without adding friction, distraction, or new safety risks?
Emergency departments are an especially demanding test environment because clinicians work under time pressure, interruptions, changing patient acuity and high cognitive load. A technically capable system may still fail to produce clinical value if clinicians stop using it when workload increases.
SHAKED is therefore important as a prospective implementation study. It shifts attention from what an LLM can do in a controlled benchmark toward how a decision-support system behaves when embedded in actual care.
What the Study Tested
The investigators conducted a DECIDE-AI stage 1 prospective evaluation at Rambam Health Care Campus in Haifa, Israel. During four weeks, 1,138 patients were analyzed across two parallel emergency-department units: an intervention wing with SHAKED available and a comparison wing following routine rotations.
SHAKED integrated with the electronic health record and used multiple large language models within a secure hospital environment. The study examined clinician adoption, patterns of use, consultation workflow, emergency-department length of stay, sampled output appropriateness and detected adverse events.
This design should not be described as a definitive patient-level randomized efficacy trial. It was an early-stage prospective clinical evaluation intended to characterize implementation, engagement and signals that can inform later randomized trials.
What the Study Found
Adoption declined rapidly
Clinical adoption fell from 68% to 30% over the four-week study. Engagement also declined as each shift progressed: the reported odds ratio for use was 0.72 per shift hour (95% CI 0.62–0.83). That pattern suggests that sustained use may be sensitive to workload and workflow demands.
Clinicians used the system selectively
Physicians were more likely to use SHAKED for radiology consultations, with an odds ratio of 2.98 (95% CI 1.58–5.63). This is relevant because decision support may prove more useful for specific tasks than as a universal layer applied to every patient encounter.
Sampled outputs were usually judged appropriate
Expert review rated 99 of 100 sampled outputs as clinically appropriate. That is encouraging, but it should not be translated into a claim of “99% diagnostic accuracy.” The sampled review assessed clinical appropriateness of outputs, not a validated diagnostic-accuracy endpoint across all encounters.
No throughput improvement was established
Emergency-department length of stay was 4.9 hours in both wings (P=0.99). In the intention-to-treat analysis, consultation cycle time showed a non-significant trend of approximately 9.4 minutes shorter with SHAKED (P=0.077). The study therefore did not establish a significant workflow benefit on these principal timing measures.
No adverse events were detected—but that is not proof of safety
The investigators reported no detected adverse events related to the system during the four-week evaluation. Given the study duration, sample size and early-stage design, this finding should be interpreted as reassuring surveillance within this pilot—not as evidence that uncommon harms or downstream clinical risks have been excluded.
Methodological Quality
Overall assessment: informative prospective implementation evidence, but not a definitive efficacy or safety trial.
The study is stronger than a retrospective benchmark because it evaluated an LLM-based system inside routine emergency care, measured actual clinician engagement, examined workflow outcomes and prospectively monitored for problems. It also followed the DECIDE-AI framework for early-stage evaluation of AI decision-support systems.
At the same time, the design limits causal claims about clinical effectiveness. The study was conducted at a single tertiary center for only four weeks. Exposure depended on clinical-unit workflow and actual physician activation rather than a simple individually randomized patient-level intervention. Per-protocol analyses may therefore be influenced by selective adoption: the cases in which clinicians chose to use the system may differ systematically from those in which they did not.
Strengths
- Prospective real-world evaluation: the system was tested during routine emergency-department care rather than only on retrospective datasets.
- Meaningful scale for an early implementation study: 1,138 patients were analyzed over four weeks.
- Human-factors measurement: adoption, workload-sensitive disengagement and task-specific use were treated as outcomes rather than ignored.
- Multiple outcome domains: the investigators evaluated workflow, clinician behavior, expert-reviewed outputs and detected adverse events.
- Transparent reproducibility resources: analysis and figure-generation code are publicly available under an MIT license, while a de-identified patient-level analytic dataset is available through controlled access.
- Explicitly cautious interpretation: the authors state that the findings inform future randomized-trial design but do not justify clinical deployment at this stage.
Limitations and Cautions
- The study was conducted at a single tertiary emergency department, limiting external generalizability.
- The evaluation lasted only four weeks, which is insufficient to establish durable adoption or long-term safety.
- Clinical adoption dropped from 68% to 30%, raising questions about sustained usability and workflow fit.
- The design was an early prospective parallel-unit implementation evaluation, not a definitive patient-level randomized efficacy trial.
- Per-protocol comparisons may be affected by selective-adoption bias.
- Exploratory comparisons reported in extended analyses used unadjusted P values for multiple comparisons, increasing the risk of chance-positive findings.
- Only 100 outputs were sampled for expert appropriateness review; 99 appropriate outputs should not be generalized to all possible clinical situations.
- No adverse events were detected, but the study was not powered to exclude uncommon or delayed harms.
- No significant improvement was demonstrated in emergency-department length of stay, and the intention-to-treat consultation-cycle result did not reach statistical significance.
Funding and Conflicts of Interest
The study received AWS Partner funding and US$20,000 in cloud-computing credits, which supported secure infrastructure, AWS Bedrock access and scalable LLM inference. According to the publication, the funder had no role in study design, data collection or analysis, the decision to publish, or manuscript preparation.
The authors declared no competing interests. The infrastructure support remains relevant context because cloud architecture was part of the technical deployment environment, but the disclosed support alone is not evidence that the findings are invalid.
Soleil Insight
The central lesson from SHAKED is that technical plausibility is not the same as clinical adoption.
A system can generate outputs that experts usually judge appropriate and still lose users rapidly during ordinary clinical work. The fall from 68% to 30% adoption may therefore be as scientifically important as any model-performance metric. It suggests that future medical-AI trials must treat workflow fit, cognitive load, trust, usability and sustained engagement as core components of effectiveness—not as secondary implementation details.
There is also an important safety distinction. “No adverse events detected” over four weeks is useful evidence, but it is not equivalent to establishing safety. Responsible clinical AI requires larger and longer prospective studies capable of detecting rare errors, downstream consequences and changes in clinician behavior.
The appropriate conclusion is neither that LLM clinical decision support has failed nor that it is ready for routine use. SHAKED provides a valuable bridge from laboratory-style AI evaluation to real-world clinical testing, while showing exactly why deployment decisions require evidence about the human-system interaction, not only the model.
Original Publication
Leibovitch L, Ahituv A, Gorenshtein A, et al. Prospective evaluation of a large language model clinical decision support system in the emergency department. Nature Medicine. Published August 19, 2026.
DOI: 10.1038/s41591-026-04601-5
ClinicalTrials.gov: NCT06902675
Reproducibility: The publication reports publicly available analysis code and a permanently archived Zenodo record (DOI 10.5281/zenodo.20736931); de-identified patient-level data are available under controlled access subject to institutional and ethics requirements.
Featured Image Attribution
Featured image: Ernie Branson / National Cancer Institute. Public-domain U.S. Government image via Wikimedia Commons. The historical photograph is used illustratively to represent computer-assisted clinical information review and does not depict the SHAKED study, its participants or its software interface.
Advancing Science. Empowering Humanity.
Soleil Research provides evidence-focused scientific education and does not replace clinical judgment, institutional policy or professional medical guidance.