
Photo Credit: Moritz Erken
Scientific Frontline: Extended "At a Glance" Summary: On-Premise Medical AI Agents
The Core Concept: A locally operated diagnostic AI system designed to support clinical decision-making while ensuring data privacy and result transparency.
Key Distinction/Mechanism: Unlike cloud-based Large Language Models (LLMs), this system runs entirely on a local infrastructure, keeping sensitive patient data within the institution's control. It utilizes two interacting AI agents (simulating a doctor and a patient) and relies on diagnostic consistency across multiple evaluations to gauge reliability, referring uncertain cases to human medical professionals.
Major Frameworks/Components:
- Selective Autonomy: The AI supports decisions but transfers uncertain cases to human experts.
- Agent Interaction: A simulated environment where an "AI doctor" questions an "AI patient," requests lab values, and formulates a diagnosis with reasoning.
- Consistency Tracking: Evaluating reliability by checking if the AI reaches the same diagnosis upon repeated assessment of the same case.
- On-Premise Infrastructure: Complete local data processing to manage data protection, model versions, and access rights.
Branch of Science: Artificial Intelligence (Clinical AI), Digital Health, Medicine, and Data Science.
Future Application: Integration into clinical workflows to provide reliable diagnostic support for human doctors, particularly in triaging complex cases, while maintaining strict European data sovereignty standards.
Why It Matters: This system addresses the two primary hurdles of AI in medicine: protecting sensitive health data and providing transparent, verifiable reliability metrics for AI-generated diagnoses, ensuring humans retain ultimate medical responsibility.
AI agents could reliably support diagnoses and clinical decision-making in the future—provided that sensitive health data are protected and clinicians can assess how reliable individual AI-generated results are. Researchers at the Else Kröner Fresenius Center (EKFZ) for Digital Health at TU Dresden and Dresden University Hospital have developed an on-premises medical AI system that addresses both challenges. Their findings have been published in the journal Nature Medicine.
Large language models are increasingly capable of handling complex medical tasks. However, their use in clinical practice faces two fundamental challenges: sensitive patient data need to remain under the control of the respective institution rather than being shared with external services, and clinicians need to be able to assess the reliability of AI-generated diagnoses. A research team led by Professor Jakob N. Kather in Dresden developed a fully on-premises diagnostic AI system and evaluated it in a simulation environment in which two AI agents interact with each other: one takes on the role of the physician, and the other that of the patient. Using this setup, the researchers investigated how reliably medical AI systems can make diagnostic decisions. The system is built upon the MIRA AI agent introduced in June 2026.
High Diagnostic Accuracy in Standardized Tests The AI agent was tested using standardized clinical cases covering a range of conditions, including appendicitis, cholecystitis, pneumonia, pulmonary embolism, and urinary tract infections. The cases were based on anonymized electronic patient data, including diagnoses, laboratory values, medications, and physical examination findings. The physician AI agent could ask follow-up questions and request clinical findings and laboratory results. Based on this information, it generated a diagnosis and an accompanying rationale. The best locally operated AI model reached the correct diagnosis in approximately 90% of cases in one benchmark and 84% in the other. For 181 randomly selected cases, physicians additionally reviewed the diagnoses. The automated evaluation and the physicians’ consensus agreed in more than 90% of cases.
Consistent Answers Are the Strongest Indicator of a Correct Diagnosis One of the study’s key questions was how to determine whether an AI-generated decision can be trusted. To address this, the researchers investigated several indicators of diagnostic correctness. The most informative was whether the AI arrived at the same diagnosis when processing the same case repeatedly. The more stable the diagnosis remained across multiple runs, the more likely it was to be correct. The model’s internal likelihood scores were less useful for predicting whether a diagnosis was correct. This was also evident in a stress test: when the researchers removed reliable information from the system, diagnostic accuracy dropped substantially. Although the AI diagnoses became more frequently incorrect, the model’s probability-based signals did not reflect this.
Medical AI Agents to Support Clinical Practice Based on their findings, the researchers propose a possible framework for collaboration between clinicians and AI. Rather than requiring clinicians to either fully trust an AI system or verify every decision, the framework would distinguish cases with stronger reliability signals from more uncertain ones. Uncertain cases would always be deferred to medical professionals for review. Professor Jakob N. Kather, senior author of the study and Professor of Clinical Artificial Intelligence at TU Dresden, explains: “Our goal is an AI agent with selective autonomy. These systems should support clinicians in decision-making but never take over completely. It is therefore essential that AI-generated results are understandable to clinicians and that the systems clearly indicate when their outputs are uncertain. Responsibility for diagnosis and treatment will always remain with humans.”
Local Deployment Enables Institutional Control A second focus of the study is technical control over the AI system. The AI models and all data processing ran entirely within locally operated infrastructure. This gives medical institutions greater control over where data are processed and which AI models are used. On-premises infrastructure provides the technical basis for managing data protection, model versions, access rights, and monitoring processes within the respective institution. Kather also sees this as a prerequisite for advancing medical AI responsibly in Europe: “Calls to slow down AI development are, in my view, heading in the wrong direction. In medicine, we are still at the very beginning: AI systems such as our agents are only just starting to show what is possible, and patients have so far seen very little benefit. We need to move faster, not slower. What matters is that we develop these systems in a way that is safe, transparent, and preserves data sovereignty—for example, by running them entirely within the hospital. But for this to work, Europe also needs to play a role in the foundational technology itself, including the language models, rather than simply building on technologies developed elsewhere.”
Next Steps Toward Clinical Use The study also highlights questions that need to be addressed before any potential clinical deployment. The researchers found indications of differences in diagnostic accuracy between patient groups. It remains unclear why simulated cases involving older patients, in particular, showed poorer results, and this will require further investigation. In addition, the findings are based on retrospective simulations using existing clinical data rather than on testing during routine clinical care. “Our results show that we have taken an important step toward more reliable medical AI agents. Next, we want to investigate how well this approach performs under realistic conditions, with clinicians involved in the process, and across a broader range of diseases and clinical data,” says Li Zhang, first author of the publication and a researcher on Professor Kather’s team. Another focus is efficiency. Because estimating reliability currently requires repeated runs for each case, the team is working to reduce the computational cost without compromising the quality of the results. In addition to the Dresden researchers, scientists from the National Center for Tumor Diseases (NCT) Heidelberg at Heidelberg University Hospital also contributed to the study.
Published in journal: Nature Medicine
Title: On-premise medical AI agents for reliable clinical decision-making
Authors: Li Zhang, Georg Wölflein, Dyke Ferber, Junhao Liang, Zunamys I. Carrero, Xuewei Wu, Julien Vibert, Jan Clusmann, Lino Möhrmann, Elena E. Möhrmann, Catharina Wichmann, Fabian Wolf, Tim Lenz, and Jakob Nikolas Kather
Source/Credit: Dresden University of Technology
Edited by: Scientific Frontline
Reference Number: ai091526_01