Software

TU Dresden: Dresden Researchers Develop Locally Operated AI Agent to Provide Reliable Support for Medical Decisions

September 15, 2026. In the future, AI agents could reliably support diagnoses and clinical decisions—provided that sensitive health data is protected and the reliability of the results is transparent and verifiable for physicians. Researchers at the Else Kröner Fresenius Center (EKFZ) for Digital Health at TU Dresden and Dresden University Hospital have now developed a locally operated AI system that meets both of these requirements. Their findings were published in the journal *Nature Medicine*.

Share this Post
Stock image: Artificial Intelligence (AI) / Photo by Shubham Dhage on Unsplash

Contact info

Silicon Saxony

Marketing, Kommunikation und Öffentlichkeitsarbeit

Manfred-von-Ardenne-Ring 20 F

Telefon: +49 351 8925 886

redaktion@silicon-saxony.de

Large language models are now capable of handling complex medical tasks. However, their use in clinical practice faces two fundamental challenges: Sensitive patient data must remain under the control of the respective institution and not be shared with external services; and physicians must be able to assess how reliable the AI’s decisions are. A research team led by Prof. Jakob N. Kather from Dresden developed a fully locally operated diagnostic AI system and tested it in a simulation environment in which two AI agents interact with each other: one takes on the role of the physician, the other that of the patient. Using this setup, the researchers investigated how reliably medical AI systems make diagnostic decisions. The system is based on the AI agent MIRA, which was introduced in June 2026.

High Diagnostic Accuracy in Standardized Tests

The AI agent was tested using standardized clinical case studies that included various conditions, such as appendicitis, cholecystitis, pneumonia, pulmonary embolism, and urinary tract infections. The tests were based on anonymized electronic patient data, such as diagnoses, lab results, medications, and examination findings. The medical AI agent was able to ask follow-up questions and request test results and laboratory values. Based on this information, it formulated a diagnosis and a rationale. The best locally operated AI model arrived at a correct diagnosis in 84 percent of cases in one test and in 90 percent of cases in another. For 181 randomly selected cases, physicians additionally reviewed the diagnoses. The automated assessment and the medical consensus agreed in over 90 percent of the cases.

Consistency as the Most Important Indicator of Correct Diagnoses

One question addressed by the study was how to determine whether an AI decision can be trusted. To this end, the researchers examined indicators of whether a diagnosis was correct. Particularly telling was whether the AI arrived at the same diagnosis when processing the same case repeatedly. The more stable the diagnosis was across multiple runs, the more often it was actually correct. The accuracy of a diagnosis was less reliably assessed based on the model’s internal probability values. This was also evident in a stress test: When the researchers deprived the system of reliable information, diagnostic accuracy dropped significantly. Although the AI diagnoses were more often incorrect, the model’s probability values did not reflect this.

Medical AI Agents as Support for Everyday Clinical Practice

Based on their findings, the researchers propose a possible framework for collaboration between medical professionals and AI: Instead of either trusting an AI system completely or having to verify every decision, cases with particularly reliable diagnoses could be distinguished from uncertain ones. The system would refer uncertain cases to medical professionals for review. Prof. Jakob N. Kather, the study’s last author and professor of Clinical Artificial Intelligence at TU Dresden, emphasizes: “Our goal is an AI agent with selective autonomy. These systems are intended to support physicians in decision-making, but never to take over completely. That is why it is important that AI results are comprehensible to physicians and that the systems transparently indicate where they are uncertain. However, responsibility for diagnosis and treatment will always remain with humans in the future.”

Local operation enables institutional control

A second focus of the work is technical control over the AI system. The AI models and data processing run entirely within a locally operated infrastructure. This allows medical institutions to exercise greater control over where data is processed and which AI models are used. The local infrastructure creates the technical conditions necessary to manage data protection, model versions, access rights, and monitoring processes within the respective institution. Kather also sees this as a prerequisite for the responsible further development of medical AI in Europe: “In my view, calls to slow down AI development are heading in the wrong direction. In medicine, we’re just getting started: AI systems, like our agents, are only beginning to show what’s possible, and patients have barely benefited from them so far. We need to pick up the pace, not slow down. It is crucial that we develop the systems in such a way that they are safe, transparent, and ensure data sovereignty—for example, by operating them entirely within the clinical setting. To achieve this, however, Europe must also play a role in the future when it comes to the underlying technology—that is, the language models themselves—rather than simply building on technologies developed by others.”

Next Steps Toward Clinical Implementation

The study also highlights which questions still need to be addressed before any potential clinical implementation. For example, the researchers found evidence of differences in diagnostic accuracy between patient groups. It remains unclear why simulated cases involving older patients in particular showed poorer results, and this will be investigated further. Additionally, the results are based on retrospective simulations using existing clinical data rather than on testing during ongoing clinical operations. “Our results show that we have taken an important step forward on the path to reliable medical AI agents. Next, we want to investigate how this approach performs under realistic conditions, with physicians involved in the process and with a broader spectrum of diseases and clinical data,” says Li Zhang, first author of the publication and a researcher on Prof. Kather’s team. Another focus is on efficiency: Since assessing reliability has so far required multiple runs per case, the team is working to reduce the computational effort in the future without compromising the quality of the results. In addition to the Dresden researchers, scientists from the National Center for Tumor Diseases (NCT) Heidelberg at Heidelberg University Hospital were also involved in the study.

Publication

Li Zhang, Georg Wölflein, Dyke Ferber, Junhao Liang, Zunamys I. Carrero, Xuewei Wu, Julien Vibert, Jan Clusmann, Lino Möhrmann, Elena E. Möhrmann, Catharina Wichmann, Fabian Wolf, Tim Lenz, Jakob Nikolas Kather: On-Premise Medical AI Agents for Reliable Clinical Decision-Making, Nature Medicine, 2026. DOI: https://www.nature.com/articles/s41591-026-04609-x

Else Kröner Fresenius Center (EKFZ) for Digital Health

The EKFZ for Digital Health at the Faculty of Medicine of the Technical University of Dresden (TUD) and the Carl Gustav Carus University Hospital in Dresden was founded in September 2019. It is funded by the Else Kröner-Fresenius Foundation with a grant of 40 million euros for a period of ten years. The center focuses its research activities on innovative medical and digital technologies at the direct interface with patients. The goal is to fully harness the potential of digitalization in medicine to sustainably improve healthcare, medical research, and clinical practice.

Contact

EKFZ for Digital Health
Dresden University of Technology
Anja Stübner and Dr. Viktoria Bosak
Science Communication
Tel.: +49 351 – 458 11379
news.ekfz@tu-dresden.de
digitalhealth.tu-dresden.de 

– – – – –

Related Links

👉 https://tu-dresden.de   

Photo: unsplash

You may be interested in the following

Contact info

Silicon Saxony

Marketing, Kommunikation und Öffentlichkeitsarbeit

Manfred-von-Ardenne-Ring 20 F

Telefon: +49 351 8925 886

redaktion@silicon-saxony.de