MoBayes

A Modular Bayesian Framework for Separating Reasoning from Language in Conversational Clinical Decision Support

  • Yusuf Kesmen1*
  • Fay Elhassan1*
  • Jiayi Ma2*
  • Julien Stalhandske1
  • Yena Chang1
  • David Sasu1
  • Alexandra Kulinkina1
  • Akhil Arora3†
  • Lars Klein1†
  • Mary-Anne Hartley1†

1EPFL  ·  2University of Bern  ·  3Aarhus University
*equal contribution  ·  †equal supervision

Abstract

Large language models (LLMs) are increasingly used for conversational clinical decision support, yet they conflate next token prediction with probabilistic decision making. We argue that this conflation reflects an architectural limitation: such systems lack explicit posterior tracking, controllable abstention thresholds, and auditable reasoning chains. We introduce MoBayes, a Modular Bayesian dialogue framework that separates reasoning from language. The LLM acts only as a language interface, parsing patient conversation into structured observations, while a Bayesian module performs probabilistic inference over these observations to update posteriors, select follow up questions via expected information gain and determine when to stop or defer through calibrated decision thresholds. This design enables explicit posterior tracking, controllable selective decision making, and replaceable population specific statistical backends without retraining the language model. Across empirical and LLM generated knowledge bases, MoBayes outperforms standalone frontier LLM doctors, including matched model family comparisons where inexpensive sensor models paired with MoBayes exceed larger autonomous models at lower cost. The advantage persists under adversarial patient communication styles and across varying diagnostic scenarios. These results suggest that reliable conversational clinical decision support systems should separate probabilistic reasoning from language generation rather than scaling model size alone. Code is available online.

Clinical statistics exist outside the model. Population statistics Observed data Disease priors · feature likelihoods TRAINING Language model Training record Patient A Test result Positive Model answer “Patient A’s test was positive.” Positive New population 5% Disease prevalence Stored prior 20% 5% Fair coin 50%50% HeadsTails Model estimate 80%20% HeadsTails
Training can bury clinical statistics in model weights. MoBayes keeps knowledge explicit and replaceable.

A clinical reasoning harness

Medical training and agentic workflows often leave diagnosis, question selection and stopping to the language model. Its diagnostic beliefs are difficult to inspect or change independently.

Other approaches introduce Bayesian inference, but still derive their probabilities from the language model’s predictions.

MoBayes separates conversation from clinical inference. The language model extracts evidence and phrases questions. A Bayesian engine updates disease probabilities, chooses questions and decides when to stop or abstain.

This separation supports robust, low cost reasoning with controllable decisions and calibrated stopping thresholds.

MoBayes workflow

MoBayes Patient Language model I have a fever. Any wheezing? Before evidence Stop at 80% 0100% 51.6% < 80% Ask again Information gain 00.1 bits Bayesian model FindingFever ValueYes CertaintyCertain

Use clinical knowledge without retraining

Keep clinical probabilities in a table that can be inspected, updated and adapted to each hospital.

From hospital records

Hospital records capture local disease frequencies and findings. These form an updatable probability table without training a language model on patient data.

The table defines the diagnostic scope. When evidence is insufficient, the system can defer to a clinician using an adjustable confidence threshold.

Without hospital data

Without local records, structured prompts extract probabilities from a language model. The table remains inspectable and replaceable, with the same control over inference and abstention.

Reasoning over this table improves diagnostic performance while preserving these controls, without requiring a larger conversational model.

Estimate probabilities from observations. ObservationsPatient recordsLanguage modelP(fever | pneumonia)? or Knowledge baseConditionFeverCoughWheezePneumonia0.800.850.15Bronchitis0.350.900.25Asthma0.100.400.80P(finding | condition) Population A priorsSame language model

Results

Better diagnostic decisions with a smaller language model.

  1. Small sensors, strong decisions.

    Replace a frontier doctor with a smaller, same-family language sensor inside MoBayes. Each arrow moves toward higher DHS and lower session cost.

    MoBayes + sensorStandalone
    Cost per consultation, including patient simulation.
    View Table 2 Sensors vs. frontier doctors

    50 cases per system and setting. MoBayes uses a DHS-optimal threshold per sensor.

    Top-1 and SA are percentages; DHS balances accuracy and coverage. This experiment is separate from Table 1. Table 2 in the paper.

  2. Less language work per case.

    The LLM only parses and verbalizes; Bayesian updates require no API calls. Knowledge-base construction is amortized across consultations.

    MoBayes + sensorStandalone
    View paper figure.
  3. Robust to the patient.

    The advantage persists across patient simulation protocols. MoBayes also stays stable when patients exaggerate, withhold or bury information.

    MoBayesStandalone
  4. Choose when to commit.

    Raise τ to trade coverage for selective accuracy. The same model can answer more cases or defer more cautiously.

    77%Selective accuracy
    96%Coverage
    86DHS
    More coverageMore selective
    Square: standalone model.

BibTeX

@misc{kesmen2026mobayes,
  title         = {MoBayes: A Modular Bayesian Framework for Separating Reasoning
                   from Language in Conversational Clinical Decision Support},
  author        = {Yusuf Kesmen and Fay Elhassan and Jiayi Ma and Julien Stalhandske
                   and Yena Chang and David Sasu and Alexandra Kulinkina
                   and Akhil Arora and Lars Klein and Mary-Anne Hartley},
  year          = {2026},
  eprint        = {2604.20022},
  archivePrefix = {arXiv},
  primaryClass  = {cs.LG},
  url           = {https://arxiv.org/abs/2604.20022}
}