Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support
06:00 · August 3, 2026 · arXiv cs.AI RSS

LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for symptom assessment and clinical decision support in diagnostic and treatment guidance, administrative documentation, and rules-based alert enhancement. This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop. For that task, the evidence of safety does not yet exist. The gap is not in medical knowledge but in the fidelity of clinical evaluation: a model optimized to continue the most probable text is not optimized to act safely when the safe answer is the improbable must-not-miss diagnosis. Safe triage is not the selection of the most likely diagnosis; it is a sequential decision under asymmetric cost, in which the single catastrophic miss outweighs many false alarms, and the decisive signal may be one the patient has not volunteered - and that the model has not been trained to seek. The core deficit is therefore one of information gathering under uncertainty. Under incomplete histories, LLM systems may fail to show the behaviors safe triage requires: broadening the differential; seeking the missing red flag; lowering the threshold for escalation; deferring judgement until sufficient information is obtained; and escalating concern where high-harm diagnoses remain unexcluded. These modes of failure for LLMs can be difficult to detect considering that evaluations to date often use complete, well-curated, confidence-gated simulations. The application of LLMs under these conditions may be amplified by assistant-like behaviors and positive bias, including credulity, agreeableness, and miscalibration - when these are not constrained by clinical triage logic.
Summary
Large language models now match or exceed physicians on medical licensing exams and curated diagnostic benchmarks, yet this performance rests on complete, clinician-structured cases rather than the incomplete, unstructured accounts typical of real patients. The Perspective identifies a fundamental mismatch: models trained to predict the most probable next token are not optimized to gather missing information or to prioritize rare but high-harm “must-not-miss” diagnoses when the safe action is the statistically improbable one. Safe triage therefore requires sequential decisions under asymmetric error costs, in which a single catastrophic miss outweighs multiple false alarms and the decisive signal may never be volunteered by the patient.
Empirical evidence underscores the gap. In one randomized study, frontier models that identified the correct condition in 94.9 percent of curated scenarios succeeded in fewer than 34.5 percent of cases once members of the public described their own problems in open dialogue; participants using the models performed no better than those receiving no assistance. Systematic reviews further show that existing benchmarks lack construct validity for this setting, concentrating model weaknesses precisely on rare and ambiguous presentations where safety matters most. Assistant-like tendencies such as credulity, agreeableness, and miscalibration can compound these deficits when they are not overridden by explicit triage logic.
The authors therefore argue that claims of readiness for autonomous clinical decision support must be supported by new evaluation regimes. These should incorporate incomplete histories, harm-weighted metrics that penalize missed high-stakes diagnoses more heavily than over-triage, and validated simulators capable of testing information-gathering behavior under uncertainty before any deployment in unsupervised patient-facing roles.
Why it matters
Directly addresses ethical, safe, and transparent AI deployment in healthcare, aligning with Dutch/EU priorities on responsible AI and high-risk medical systems under the AI Act. Offers actionable insights for researchers evaluating clinical LLMs.






