A medical chatbot can perform well when researchers give it a complete case. A person must first decide what to tell it, make sense of its reply and choose what to do. A 2026 paper in Nature Medicine puts that whole exchange under the microscope. The experiment used written scenarios, not treatment of real patients.

THE SHORT READ
  • The important comparison is people with AI versus people using their usual resources.
  • Recognizing a possible condition and choosing an appropriate next step are different outcomes.
  • Publication in 2026 does not make the tested 2024 model versions a ranking of today’s products.
THE GUIDE AT A GLANCEBean et al. · Human–AI interaction
Publication
9 February 2026 · Nature Medicine
Participants
1,298 UK adults; 10 written scenarios
Comparison
Three chatbot groups versus usual information sources
Data collection
21 August–14 October 2024

Original publication: 9 Feb 2026 · The date above refers to this brief.

THE NUMBERS, IN CONTEXT

The models alone: choosing the expected next step

Correct responses (%) · direct prompts with the complete scenario

GPT-4o64.7
Llama 348.8
Command R+55.5
0100
Source: Bean et al., task-validation experiment; 60 responses per model per scenario across 10 scenarios. These are model-only point estimates, not participant results or treatment effects. Uncertainty intervals are not shown. The randomized human comparison did not find a significant improvement in next-step accuracy.

What the researchers actually tested

Bean and colleagues randomly assigned participants to use GPT-4o, Llama 3, Command R+ or their usual resources. Each person could complete up to two scenarios. The researchers collected 600 scenario responses per arm: these are responses, not 600 different people in each group. Participants named possible conditions and selected a next step on a five-option scale. Doctors supplied the reference answers.

The study was preregistered, but technical problems led to participant replacements and a change in the stopping rule. The paper describes these departures. The model-only comparisons were added after preregistration, so they should not be confused with the original randomized comparison.

Source 1 ↗

The result that matters for the headline

Participants using the chatbots were worse than the control group at naming at least one relevant condition. The control group’s odds were 1.76 times those of the combined chatbot groups (95% confidence interval 1.45–2.13). This is an odds ratio, not a percentage-point difference or a personal-risk estimate.

For selecting the appropriate next step, none of the chatbot groups differed significantly from control. That does not establish exact equality. Separately, the models did much better at identifying conditions when given the full scenario directly. A model’s answer sheet and a person’s final decision were measuring different things.

Source 1 ↗

Why a fluent answer can lose its value in the conversation

The authors inspected interactions and identified incomplete information supplied by users, misinterpretation by models and relevant suggestions that did not appear in the person’s final answer. These observations help explain possible failure points; they do not tell us what percentage of the total effect each mechanism caused.

Our reading is that the interface belongs inside the experiment. A product that asks a useful follow-up question changes the information available. A product that lists many possibilities changes the reader’s task. Evaluating only the final model response leaves both decisions outside the score.

Source 1 ↗

How this connects to the other two studies

The linked image study asks whether accompanying words can pull a model away from visual evidence. The linked education study asks what an explanation does to a novice’s answer. Together they motivate three separate checks: the input, the model output and the human response. They are different experiments with different populations and endpoints, not three estimates to pool into one “AI accuracy” percentage.

For your reading notes, write down exactly who or what received a score. Then record what information was supplied and what comparison would be needed for the claim being made.

Source 2 ↗Source 3 ↗

What would make the next study more useful?

Our proposed next test would compare a fixed chatbot interface with one designed to ask for missing information, using the same model and a preregistered scoring plan. It would recruit people with different levels of reading and digital confidence, report incomplete sessions, and test whether gains persist across scenarios. Those are design proposals, not demonstrated improvements.

This study cannot establish how newer systems perform, how people behave during a real emergency, or whether a tool improves health outcomes. Its useful contribution is a way to ask a harder question: can the person using the system reach a better decision?

Source 1 ↗
CONNECT THE EVIDENCE

Three questions hidden inside “Does it work?”

TestWhat is scoredWhat remains outside it
Model aloneAnswer to a fully supplied caseWhat the user omits or misunderstands
Person with a chatbotThe person’s final answerActual care received and health outcomes
Prospective care studyA predefined outcome in a real workflowOther settings and future software versions

Our reading framework. The third row describes a different study design, not an experiment reported in this paper.

READER QUESTIONS

Your questions, answered

Did the researchers test real patients receiving treatment?

No. Adults worked through written scenarios. The study evaluates decisions in that task; it does not measure recovery, hospital admissions or harms from actual treatment.

Does “no significant difference” prove chatbots have no effect?

No. It means the comparison did not establish a difference under the study’s analysis. Precision, the measured outcome and the tested conditions still matter.

Can I use these results to rank current chatbots?

No. The human experiment took place in 2024 and used particular models and an interface. A current product needs its own evaluation.

What is the most useful question to ask of an AI health headline?

Who did better: the model answering a question, a person using it, or a patient receiving care? Ask for the comparison behind that statement.

LIMITATIONS

Limits of this interpretation

  • Ten scenarios and a UK, English-speaking participant sample limit generalization.
  • Participant replacements and the changed stopping rule should be considered alongside the preregistration.
  • A controlled scenario task cannot establish patient benefit or the performance of later model versions.
SOURCE NOTES

Sources & transparency

  1. Bean, Payne, Parsons et al. (2026): Reliability of LLMs as medical assistants for the general public

    Current publisher HTML: abstract, results, discussion and methods checked. Supplementary files and raw data were not independently assessed. The page links a 17 April 2026 correction; the separate correction text could not be retrieved. Our wording is original; this source is CC BY 4.0. · Accessed 28 Sep 2026

    DOI: 10.1038/s41591-025-04074-y
  2. Buckley, Diao, Srivastava et al. (2026): Multimodal foundation models exploit text to make medical image predictions

    Publisher-indexed results and figure 2 caption, plus the primary paper’s PMC abstract and publication record, checked. Direct publisher retrieval failed; supplements, code and raw data were not checked. Only a short factual account is used; journal figures are not reproduced. · Accessed 28 Sep 2026

    DOI: 10.1038/s41467-026-74207-5
  3. Teng, Tan, Cao et al. (2026): Impact of AI misinformation on diagnostic accuracy and confidence calibration in novice medical students

    Publisher-indexed abstract, results and study-design sections checked. Direct full-page retrieval failed; supplements, participant data and registration record were not independently assessed. Only a short factual account is used. · Accessed 28 Sep 2026

    DOI: 10.1038/s41746-026-02547-z
  4. Creative Commons Attribution 4.0 — licence for the Bean et al. paper

    Licence linked by the publisher. Our explanation and chart are newly written, with attribution to the original authors; no publisher figure is reproduced. · Accessed 28 Sep 2026

Prepared and source-checked with AI on 28 September 2026 (UTC). Press-news Team is our collective publication byline, not a claim of medical credentials or human review. No human editorial or clinical review has taken place. This is an educational account of how AI was evaluated, not medical advice or a clinical endorsement. Access limitations are listed for each source; proposed follow-up tests are our commentary, not reported findings.

Source check: AI source check — selected primary research and reported numbers

Clinical review: Not applicable to this educational guide

Suggest a correction