A medical chatbot can perform well when researchers give it a complete case. A person must first decide what to tell it, make sense of its reply and choose what to do. A 2026 paper in Nature Medicine puts that whole exchange under the microscope. The experiment used written scenarios, not treatment of real patients.
- The important comparison is people with AI versus people using their usual resources.
- Recognizing a possible condition and choosing an appropriate next step are different outcomes.
- Publication in 2026 does not make the tested 2024 model versions a ranking of today’s products.
- Publication
- 9 February 2026 · Nature Medicine
- Participants
- 1,298 UK adults; 10 written scenarios
- Comparison
- Three chatbot groups versus usual information sources
- Data collection
- 21 August–14 October 2024
Original publication: 9 Feb 2026 · The date above refers to this brief.
The models alone: choosing the expected next step
Correct responses (%) · direct prompts with the complete scenario
What the researchers actually tested
Bean and colleagues randomly assigned participants to use GPT-4o, Llama 3, Command R+ or their usual resources. Each person could complete up to two scenarios. The researchers collected 600 scenario responses per arm: these are responses, not 600 different people in each group. Participants named possible conditions and selected a next step on a five-option scale. Doctors supplied the reference answers.
The study was preregistered, but technical problems led to participant replacements and a change in the stopping rule. The paper describes these departures. The model-only comparisons were added after preregistration, so they should not be confused with the original randomized comparison.
Source 1 ↗The result that matters for the headline
Participants using the chatbots were worse than the control group at naming at least one relevant condition. The control group’s odds were 1.76 times those of the combined chatbot groups (95% confidence interval 1.45–2.13). This is an odds ratio, not a percentage-point difference or a personal-risk estimate.
For selecting the appropriate next step, none of the chatbot groups differed significantly from control. That does not establish exact equality. Separately, the models did much better at identifying conditions when given the full scenario directly. A model’s answer sheet and a person’s final decision were measuring different things.
Source 1 ↗Why a fluent answer can lose its value in the conversation
The authors inspected interactions and identified incomplete information supplied by users, misinterpretation by models and relevant suggestions that did not appear in the person’s final answer. These observations help explain possible failure points; they do not tell us what percentage of the total effect each mechanism caused.
Our reading is that the interface belongs inside the experiment. A product that asks a useful follow-up question changes the information available. A product that lists many possibilities changes the reader’s task. Evaluating only the final model response leaves both decisions outside the score.
Source 1 ↗How this connects to the other two studies
The linked image study asks whether accompanying words can pull a model away from visual evidence. The linked education study asks what an explanation does to a novice’s answer. Together they motivate three separate checks: the input, the model output and the human response. They are different experiments with different populations and endpoints, not three estimates to pool into one “AI accuracy” percentage.
For your reading notes, write down exactly who or what received a score. Then record what information was supplied and what comparison would be needed for the claim being made.
Source 2 ↗Source 3 ↗What would make the next study more useful?
Our proposed next test would compare a fixed chatbot interface with one designed to ask for missing information, using the same model and a preregistered scoring plan. It would recruit people with different levels of reading and digital confidence, report incomplete sessions, and test whether gains persist across scenarios. Those are design proposals, not demonstrated improvements.
This study cannot establish how newer systems perform, how people behave during a real emergency, or whether a tool improves health outcomes. Its useful contribution is a way to ask a harder question: can the person using the system reach a better decision?
Source 1 ↗Three questions hidden inside “Does it work?”
| Test | What is scored | What remains outside it |
|---|---|---|
| Model alone | Answer to a fully supplied case | What the user omits or misunderstands |
| Person with a chatbot | The person’s final answer | Actual care received and health outcomes |
| Prospective care study | A predefined outcome in a real workflow | Other settings and future software versions |
Our reading framework. The third row describes a different study design, not an experiment reported in this paper.
Your questions, answered
Did the researchers test real patients receiving treatment?
No. Adults worked through written scenarios. The study evaluates decisions in that task; it does not measure recovery, hospital admissions or harms from actual treatment.
Does “no significant difference” prove chatbots have no effect?
No. It means the comparison did not establish a difference under the study’s analysis. Precision, the measured outcome and the tested conditions still matter.
Can I use these results to rank current chatbots?
No. The human experiment took place in 2024 and used particular models and an interface. A current product needs its own evaluation.
What is the most useful question to ask of an AI health headline?
Who did better: the model answering a question, a person using it, or a patient receiving care? Ask for the comparison behind that statement.
Limits of this interpretation
- Ten scenarios and a UK, English-speaking participant sample limit generalization.
- Participant replacements and the changed stopping rule should be considered alongside the preregistration.
- A controlled scenario task cannot establish patient benefit or the performance of later model versions.
Sources & transparency
- Bean, Payne, Parsons et al. (2026): Reliability of LLMs as medical assistants for the general public
Current publisher HTML: abstract, results, discussion and methods checked. Supplementary files and raw data were not independently assessed. The page links a 17 April 2026 correction; the separate correction text could not be retrieved. Our wording is original; this source is CC BY 4.0. · Accessed 28 Sep 2026
DOI: 10.1038/s41591-025-04074-y - Buckley, Diao, Srivastava et al. (2026): Multimodal foundation models exploit text to make medical image predictions
Publisher-indexed results and figure 2 caption, plus the primary paper’s PMC abstract and publication record, checked. Direct publisher retrieval failed; supplements, code and raw data were not checked. Only a short factual account is used; journal figures are not reproduced. · Accessed 28 Sep 2026
DOI: 10.1038/s41467-026-74207-5 - Teng, Tan, Cao et al. (2026): Impact of AI misinformation on diagnostic accuracy and confidence calibration in novice medical students
Publisher-indexed abstract, results and study-design sections checked. Direct full-page retrieval failed; supplements, participant data and registration record were not independently assessed. Only a short factual account is used. · Accessed 28 Sep 2026
DOI: 10.1038/s41746-026-02547-z - Creative Commons Attribution 4.0 — licence for the Bean et al. paper
Licence linked by the publisher. Our explanation and chart are newly written, with attribution to the original authors; no publisher figure is reproduced. · Accessed 28 Sep 2026
Prepared and source-checked with AI on 28 September 2026 (UTC). Press-news Team is our collective publication byline, not a claim of medical credentials or human review. No human editorial or clinical review has taken place. This is an educational account of how AI was evaluated, not medical advice or a clinical endorsement. Access limitations are listed for each source; proposed follow-up tests are our commentary, not reported findings.
Source check: AI source check — selected primary research and reported numbers
Clinical review: Not applicable to this educational guide
Suggest a correction