A medical model can become better at answering a question while the hardest question remains unanswered: does the whole system help someone make a better decision? A paper published today offers a useful starting point for that distinction.

THE SHORT READ
  • Name the model variant, task and comparator before repeating a score.
  • A response rating, a simulated encounter and a patient outcome measure different things.
  • Our proposed next test measures the workflow, including error detection and correction.
THE STUDY AT A GLANCEMedGemma · Sellergren et al.
Publication
Nature Medicine · 6 October 2026
Research stage
Model evaluation, not a patient-outcome trial
Focus here
Factually inaccurate response ratings

Original publication: 6 Oct 2026 · The date above refers to this brief.

THE NUMBERS, IN CONTEXT

Responses rated factually inaccurate

percent of evaluated responses

Gemma 3 4B18.5
MedGemma 4B15.4
025
Reported percentages in the paper’s open-ended answer evaluation. Three physician raters per question are reported; we have not extracted rating-level denominators or confidence intervals. Newly drawn bars, not a patient-outcome comparison. Difference: 18.5 − 15.4 = 3.1 percentage points.

What was reported today

Sellergren and colleagues describe MedGemma, a family based on Gemma 3. In an open-ended answer evaluation, the 4B variant had 15.4% of responses rated factually inaccurate, versus 18.5% for Gemma 3 4B. The main text reports three physician raters per question. The smaller models also struggled with instructions in the simulated-agent task. These are distinct evaluations, not one overall clinical accuracy score.

Source 1 ↗

Put the error beside the improvement

Subtracting the two reported percentages gives 3.1 percentage points. That is a change in rated responses. It cannot be converted into the number of patients helped, because a response is not a patient and an expert rating is not a measured care outcome.

Our chart deliberately keeps the remaining inaccurate-response fraction visible. When an article shows only the improvement, ask to see the starting value, the ending value and what counted as an error. A useful model comparison keeps all three together.

Separate the model from the service

Imagine an educational tool that retrieves an image, drafts an explanation and asks a professional to approve it. The retrieval, wording, approval screen and final decision can each fail independently. Improving the generator does not automatically fix the other steps.

Our proposed evaluation would record incorrect outputs caught before use, incorrect outputs accepted, correct outputs unnecessarily rejected and time spent repairing answers. These are suggested measurements, not results reported in this paper. They make the intended service concrete enough to test.

Connect this to the human-user study

The earlier Bean study tested people working through medical scenarios, alongside a separate model-only evaluation. Its unit of interest was the interaction, rather than just an answer produced by a model. It did not test MedGemma.

Read the two papers as different questions: what can a foundation model produce, and what happens when a person uses an assistant? Our interpretation is that both questions belong in an evaluation plan. Their scores should not be placed on a single leaderboard.

Source 2 ↗

The next useful experiment

For a future deployment study, we would first specify who uses the system, what decision it supports and which mistakes matter most. Then freeze the model, prompts and interface before evaluating cases from a different setting. Report the comparison service and the cost of correction as well as answer quality.

A promising direction is easier adaptation to a carefully defined task. Demonstrating that promise would require testing the adapted system rather than assuming the foundation-model result travels unchanged. That is our research agenda, not a claim that a safe clinical service has already been established.

TRACE THE EVIDENCE

A response score is one step in a much longer workflow

Publisher main text checked; the response-rating supplement and raw ratings were not independently assessed. The older human-user paper is a design comparison, not a MedGemma test.

01What does the plotted percentage describe?

What was observed
Responses rated factually inaccurate.
Where the conclusion stops
Do not call it the fraction of patients misdiagnosed.

Source 1 · Results → Open-ended medical question answering and clinical reasoning

02Does a simulated agent establish clinical autonomy?

What was observed
The smaller models had instruction-following difficulty.
Where the conclusion stops
A simulated encounter does not test all the consequences of using an assistant with real people.

Source 1 · Results → Medical agentic behavior

Numbers you can inspect

MeasureValue & unitOrigin & method
Gemma 3 4B inaccurate-response rating18.5 percentReported Main-text percentage; rating-level data not extracted.
Source 1 · Results → Open-ended medical question answering and clinical reasoning
MedGemma 4B inaccurate-response rating15.4 percentReported Same reported comparison.
Source 1 · Results → Open-ended medical question answering and clinical reasoning
Difference in reported percentages3.1 percentage pointsCalculated 18.5 − 15.4 = 3.1; no uncertainty interval calculated.
Source 1 · Same response-rating paragraph

Compare the actual experiments

These studies answer different questions. Read the unit and endpoint before comparing results.

StudyUnit & settingReadoutInterpretation boundary
Sellergren et al. 2026

Source 1 · Results → Open-ended medical question answering and clinical reasoning

Model responsesTask and expert-rating measuresAn output evaluation is not a clinical-outcome trial.
Bean et al. 2026

Source 2 · Abstract; Figure 1

People using other assistants in scenariosConditions and proposed actionsDifferent models and units prevent a direct performance ranking.
Download evidence table (CSV)

The export includes claims, available numbers, methods and source locations. It contains our reading notes and published summaries; it is not raw participant data or an independent reanalysis.

Evidence update · 6 Oct 2026
First publication. Reported values retain their original units and source locations. Proposed next experiments are our interpretation; participant data and source-data files were not reanalysed. AI source check; no human editorial or clinical review.

CONNECT THE EVIDENCE

An evaluation ladder we would use

TestWhat to recordWhat it cannot establish alone
Answer qualityIncorrect content and uncertaintyWhether a user recognizes an error
Human interactionAccepted mistakes and correctionsAn improvement in health outcomes
DeploymentOutcomes, burden and harms against a defined comparatorPerformance in every other setting

Press-News interpretation: a proposed evaluation framework, not three completed MedGemma studies.

READER QUESTIONS

Your questions, answered

Does this mean AI is better than a doctor?

No general comparison follows from these findings. A defined benchmark or simulated task does not represent all the work involved in clinical care.

What is the difference in the chart?

3.1 percentage points, calculated as 18.5 minus 15.4. It describes the reported response-rating measure, not a patient risk reduction.

Was MedGemma tested in the earlier chatbot study?

No. We connect the studies to contrast evaluation designs, not to assign MedGemma the performance of another assistant.

What would make the next paper more useful?

A fixed, clearly described system tested with its intended users, an explicit comparison service and transparent reporting of consequential mistakes.

LIMITATIONS

Limits of this interpretation

  • Rated answers do not establish patient benefit or a safe autonomous service.
  • Percentages were taken from the main text; rating records, uncertainty estimates and Supplementary Figure 1 were not independently assessed.
  • The older human-user study tested other models and cannot establish MedGemma interaction performance.
SOURCE NOTES

Sources & transparency

  1. Sellergren, Kazemzadeh, Mahvar et al. (2026): An open vision-language model for diverse medical applications

    Publisher HTML abstract, model description, open-ended question-answering results, simulated-agent evaluation and Discussion checked. Inaccuracy percentages come from the main text description of Supplementary Figure 1; the supplement, rating records and model code were not independently inspected. Article licensed CC BY 4.0. Newly written interpretation and chart, no publisher figure reproduced. · Accessed 6 Oct 2026

    DOI: 10.1038/s41591-026-04626-w
  2. Bean, Payne, Parsons et al. (2026): Reliability of LLMs as medical assistants for the general public

    Current publisher abstract and design overview checked for the human-user versus model-only comparison. Used to explain different evaluation units; no numerical ranking against MedGemma. Raw records and supplements not reassessed. · Accessed 6 Oct 2026

    DOI: 10.1038/s41591-025-04074-y

Prepared and source-checked with AI. Press-news Team is the collective publication byline, not a medical reviewer. No human editorial or clinical review has taken place. This is an educational explanation of research methods and basic research, not an individual diagnosis or treatment recommendation. We did not conduct these experiments or reanalyse participant data. The study findings, our interpretation and suggested follow-up tests are distinguished. Source access is recorded below. Photographs are illustrative.

Source check: AI source check — linked primary passages, endpoint boundaries and printed arithmetic

Clinical review: Not applicable to this educational guide

Suggest a correction