A medical model can become better at answering a question while the hardest question remains unanswered: does the whole system help someone make a better decision? A paper published today offers a useful starting point for that distinction.
- Name the model variant, task and comparator before repeating a score.
- A response rating, a simulated encounter and a patient outcome measure different things.
- Our proposed next test measures the workflow, including error detection and correction.
- Publication
- Nature Medicine · 6 October 2026
- Research stage
- Model evaluation, not a patient-outcome trial
- Focus here
- Factually inaccurate response ratings
Original publication: 6 Oct 2026 · The date above refers to this brief.
Responses rated factually inaccurate
percent of evaluated responses
What was reported today
Sellergren and colleagues describe MedGemma, a family based on Gemma 3. In an open-ended answer evaluation, the 4B variant had 15.4% of responses rated factually inaccurate, versus 18.5% for Gemma 3 4B. The main text reports three physician raters per question. The smaller models also struggled with instructions in the simulated-agent task. These are distinct evaluations, not one overall clinical accuracy score.
Source 1 ↗Put the error beside the improvement
Subtracting the two reported percentages gives 3.1 percentage points. That is a change in rated responses. It cannot be converted into the number of patients helped, because a response is not a patient and an expert rating is not a measured care outcome.
Our chart deliberately keeps the remaining inaccurate-response fraction visible. When an article shows only the improvement, ask to see the starting value, the ending value and what counted as an error. A useful model comparison keeps all three together.
Separate the model from the service
Imagine an educational tool that retrieves an image, drafts an explanation and asks a professional to approve it. The retrieval, wording, approval screen and final decision can each fail independently. Improving the generator does not automatically fix the other steps.
Our proposed evaluation would record incorrect outputs caught before use, incorrect outputs accepted, correct outputs unnecessarily rejected and time spent repairing answers. These are suggested measurements, not results reported in this paper. They make the intended service concrete enough to test.
Connect this to the human-user study
The earlier Bean study tested people working through medical scenarios, alongside a separate model-only evaluation. Its unit of interest was the interaction, rather than just an answer produced by a model. It did not test MedGemma.
Read the two papers as different questions: what can a foundation model produce, and what happens when a person uses an assistant? Our interpretation is that both questions belong in an evaluation plan. Their scores should not be placed on a single leaderboard.
Source 2 ↗The next useful experiment
For a future deployment study, we would first specify who uses the system, what decision it supports and which mistakes matter most. Then freeze the model, prompts and interface before evaluating cases from a different setting. Report the comparison service and the cost of correction as well as answer quality.
A promising direction is easier adaptation to a carefully defined task. Demonstrating that promise would require testing the adapted system rather than assuming the foundation-model result travels unchanged. That is our research agenda, not a claim that a safe clinical service has already been established.
A response score is one step in a much longer workflow
Publisher main text checked; the response-rating supplement and raw ratings were not independently assessed. The older human-user paper is a design comparison, not a MedGemma test.
01What does the plotted percentage describe?
- What was observed
- Responses rated factually inaccurate.
- Where the conclusion stops
- Do not call it the fraction of patients misdiagnosed.
Source 1 · Results → Open-ended medical question answering and clinical reasoning
02Does a simulated agent establish clinical autonomy?
- What was observed
- The smaller models had instruction-following difficulty.
- Where the conclusion stops
- A simulated encounter does not test all the consequences of using an assistant with real people.
Source 1 · Results → Medical agentic behavior
Numbers you can inspect
| Measure | Value & unit | Origin & method |
|---|---|---|
| Gemma 3 4B inaccurate-response rating | 18.5 percent | Reported Main-text percentage; rating-level data not extracted. Source 1 · Results → Open-ended medical question answering and clinical reasoning |
| MedGemma 4B inaccurate-response rating | 15.4 percent | Reported Same reported comparison. Source 1 · Results → Open-ended medical question answering and clinical reasoning |
| Difference in reported percentages | 3.1 percentage points | Calculated 18.5 − 15.4 = 3.1; no uncertainty interval calculated. Source 1 · Same response-rating paragraph |
Compare the actual experiments
These studies answer different questions. Read the unit and endpoint before comparing results.
| Study | Unit & setting | Readout | Interpretation boundary |
|---|---|---|---|
| Sellergren et al. 2026 Source 1 · Results → Open-ended medical question answering and clinical reasoning | Model responses | Task and expert-rating measures | An output evaluation is not a clinical-outcome trial. |
| Bean et al. 2026 Source 2 · Abstract; Figure 1 | People using other assistants in scenarios | Conditions and proposed actions | Different models and units prevent a direct performance ranking. |
The export includes claims, available numbers, methods and source locations. It contains our reading notes and published summaries; it is not raw participant data or an independent reanalysis.
Evidence update · 6 Oct 2026
First publication. Reported values retain their original units and source locations. Proposed next experiments are our interpretation; participant data and source-data files were not reanalysed. AI source check; no human editorial or clinical review.
An evaluation ladder we would use
| Test | What to record | What it cannot establish alone |
|---|---|---|
| Answer quality | Incorrect content and uncertainty | Whether a user recognizes an error |
| Human interaction | Accepted mistakes and corrections | An improvement in health outcomes |
| Deployment | Outcomes, burden and harms against a defined comparator | Performance in every other setting |
Press-News interpretation: a proposed evaluation framework, not three completed MedGemma studies.
Your questions, answered
Does this mean AI is better than a doctor?
No general comparison follows from these findings. A defined benchmark or simulated task does not represent all the work involved in clinical care.
What is the difference in the chart?
3.1 percentage points, calculated as 18.5 minus 15.4. It describes the reported response-rating measure, not a patient risk reduction.
Was MedGemma tested in the earlier chatbot study?
No. We connect the studies to contrast evaluation designs, not to assign MedGemma the performance of another assistant.
What would make the next paper more useful?
A fixed, clearly described system tested with its intended users, an explicit comparison service and transparent reporting of consequential mistakes.
Limits of this interpretation
- Rated answers do not establish patient benefit or a safe autonomous service.
- Percentages were taken from the main text; rating records, uncertainty estimates and Supplementary Figure 1 were not independently assessed.
- The older human-user study tested other models and cannot establish MedGemma interaction performance.
Sources & transparency
- Sellergren, Kazemzadeh, Mahvar et al. (2026): An open vision-language model for diverse medical applications
Publisher HTML abstract, model description, open-ended question-answering results, simulated-agent evaluation and Discussion checked. Inaccuracy percentages come from the main text description of Supplementary Figure 1; the supplement, rating records and model code were not independently inspected. Article licensed CC BY 4.0. Newly written interpretation and chart, no publisher figure reproduced. · Accessed 6 Oct 2026
DOI: 10.1038/s41591-026-04626-w - Bean, Payne, Parsons et al. (2026): Reliability of LLMs as medical assistants for the general public
Current publisher abstract and design overview checked for the human-user versus model-only comparison. Used to explain different evaluation units; no numerical ranking against MedGemma. Raw records and supplements not reassessed. · Accessed 6 Oct 2026
DOI: 10.1038/s41591-025-04074-y
Prepared and source-checked with AI. Press-news Team is the collective publication byline, not a medical reviewer. No human editorial or clinical review has taken place. This is an educational explanation of research methods and basic research, not an individual diagnosis or treatment recommendation. We did not conduct these experiments or reanalyse participant data. The study findings, our interpretation and suggested follow-up tests are distinguished. Source access is recorded below. Photographs are illustrative.
Source check: AI source check — linked primary passages, endpoint boundaries and printed arithmetic
Clinical review: Not applicable to this educational guide
Suggest a correction