MoBayes
A Modular Bayesian Framework for Separating Reasoning from Language in Conversational Clinical Decision Support
1EPFL · 2University of Bern · 3Aarhus University
*equal contribution · †equal supervision
Abstract
Large language models (LLMs) are increasingly used for conversational clinical decision support, yet they conflate next token prediction with probabilistic decision making. We argue that this conflation reflects an architectural limitation: such systems lack explicit posterior tracking, controllable abstention thresholds, and auditable reasoning chains. We introduce MoBayes, a Modular Bayesian dialogue framework that separates reasoning from language. The LLM acts only as a language interface, parsing patient conversation into structured observations, while a Bayesian module performs probabilistic inference over these observations to update posteriors, select follow up questions via expected information gain and determine when to stop or defer through calibrated decision thresholds. This design enables explicit posterior tracking, controllable selective decision making, and replaceable population specific statistical backends without retraining the language model. Across empirical and LLM generated knowledge bases, MoBayes outperforms standalone frontier LLM doctors, including matched model family comparisons where inexpensive sensor models paired with MoBayes exceed larger autonomous models at lower cost. The advantage persists under adversarial patient communication styles and across varying diagnostic scenarios. These results suggest that reliable conversational clinical decision support systems should separate probabilistic reasoning from language generation rather than scaling model size alone. Code is available online.
A clinical reasoning harness
Medical training and agentic workflows often leave diagnosis, question selection and stopping to the language model. Its diagnostic beliefs are difficult to inspect or change independently.
Other approaches introduce Bayesian inference, but still derive their probabilities from the language model’s predictions.
MoBayes separates conversation from clinical inference. The language model extracts evidence and phrases questions. A Bayesian engine updates disease probabilities, chooses questions and decides when to stop or abstain.
This separation supports robust, low cost reasoning with controllable decisions and calibrated stopping thresholds.
MoBayes workflow
Use clinical knowledge without retraining
Keep clinical probabilities in a table that can be inspected, updated and adapted to each hospital.
From hospital records
Hospital records capture local disease frequencies and findings. These form an updatable probability table without training a language model on patient data.
The table defines the diagnostic scope. When evidence is insufficient, the system can defer to a clinician using an adjustable confidence threshold.
Without hospital data
Without local records, structured prompts extract probabilities from a language model. The table remains inspectable and replaceable, with the same control over inference and abstention.
Reasoning over this table improves diagnostic performance while preserving these controls, without requiring a larger conversational model.
Results
Better diagnostic decisions with a smaller language model.
-
Better diagnostic performance
16 points ahead of the strongest baseline.
90 DHSMoBayes74 DHSBest baselineFull results
100 cases per system and setting. GPT-5.4-nano backbone, except DiagnosisGPT-34B.
Top-1: correct diagnoses across all cases. SA: accuracy among committed cases. DHS: harmonic mean of SA and coverage. * Paired bootstrap p < 0.05 over the strongest baseline. Table 1 in the paper.
-
Small sensors, strong decisions.
Replace a frontier doctor with a smaller, same-family language sensor inside MoBayes. Each arrow moves toward higher DHS and lower session cost.
MoBayes + sensorStandaloneCost per consultation, including patient simulation. View Table 2 Sensors vs. frontier doctors
50 cases per system and setting. MoBayes uses a DHS-optimal threshold per sensor.
Top-1 and SA are percentages; DHS balances accuracy and coverage. This experiment is separate from Table 1. Table 2 in the paper.
-
Less language work per case.
The LLM only parses and verbalizes; Bayesian updates require no API calls. Knowledge-base construction is amortized across consultations.
MoBayes + sensorStandaloneView paper figure. -
Robust to the patient.
The advantage persists across patient simulation protocols. MoBayes also stays stable when patients exaggerate, withhold or bury information.
MoBayesStandalone -
Choose when to commit.
Raise τ to trade coverage for selective accuracy. The same model can answer more cases or defer more cautiously.
77%Selective accuracy96%Coverage86DHSMore coverageMore selectiveSquare: standalone model.
BibTeX
@misc{kesmen2026mobayes,
title = {MoBayes: A Modular Bayesian Framework for Separating Reasoning
from Language in Conversational Clinical Decision Support},
author = {Yusuf Kesmen and Fay Elhassan and Jiayi Ma and Julien Stalhandske
and Yena Chang and David Sasu and Alexandra Kulinkina
and Akhil Arora and Lars Klein and Mary-Anne Hartley},
year = {2026},
eprint = {2604.20022},
archivePrefix = {arXiv},
primaryClass = {cs.LG},
url = {https://arxiv.org/abs/2604.20022}
}