When a system accepts both pictures and words, a correct answer does not reveal which input did the work. A useful evaluation changes the inputs separately. That turns an impressive score into a question about how the answer was reached.
- Paper
- Nature Communications · 12 June 2026
- Scope
- Offline evaluation, not patient care
Original publication: 12 Jun 2026 · The date above refers to this brief.
o3 in the selected challenge set
Accuracy (%)
A short account of the finding
Buckley and colleagues evaluated eight models using 1,090 medical cases. In a selected set of 348 image questions that GPT-4V Turbo had answered correctly without accompanying clinical text, they introduced misleading context. For o3, accuracy was 84% with images alone and 28% with misleading vignettes. The misleading-text experiment used four incorrect options per question, producing 1,392 responses. These are repeated model responses, not independent patients.
Source 1 ↗Read the denominator before the bar chart
A selected subset answers a conditional question. Imagine a completely hypothetical spelling tool: researchers first keep only words it spells correctly, then add distracting hints. Its initial score in that selected sample is guaranteed to be high. The follow-up test can reveal fragility, but it cannot estimate the tool’s error rate across ordinary writing.
The same reading discipline applies to any challenge set. Look for the rule that admitted a case, whether that rule favored one system, and whether several outputs came from the same item. A thousand generated answers may contain much less variety than a thousand independently sampled problems.
A better comparison separates the inputs
Our evaluation worksheet below asks for four conditions. Image-only checks what the visual input supports. Text-only checks whether words already disclose the answer. Combined input tests integration. Conflicting input asks what the system does when the two disagree.
This is a suggested audit structure, not a promise that every study needs an identical protocol. In a real application, some text is necessary and helpful. The concern is whether the claimed capability depends on information that will be absent, unreliable or different where the system is used.
The connection to people using chatbots
Bean and colleagues studied a different transition: from a chatbot’s response to a person’s final answer. Their randomized scenario experiment found that strong standalone model performance did not translate into better next-step choices by users. That result does not validate an image model, but it shows why a chain of separately good-looking components may need evaluation as a whole.
Source 2 ↗What to watch for next
Our proposed follow-up would use a held-out collection assembled for the intended task, specify how conflicting inputs are created, and show results by case as well as by generated answer. Investigators could compare an answer-only interface with one that flags contradictions and requests clarification.
Whether those changes improve reliability is an open empirical question. The useful ambition is an evaluation that reveals when a system should hesitate, alongside when it can answer.
Our worksheet for a multimodal claim
| Input condition | Question to ask |
|---|---|
| Image only | Can the visual input support the answer? |
| Text only | Do the words already give it away? |
| Image and text | Does combining them add useful information? |
| Conflicting inputs | Does the system notice the disagreement? |
Original reading aid; these rows are questions, not additional study results.
Your questions, answered
Is a benchmark percentage my chance of receiving a correct diagnosis?
No. A score has a defined item set and scoring rule. Turning it into a personal probability would require evidence about the actual population, workflow and decision.
Is using the text always a problem?
No. Context can be essential. A credible evaluation should show when it helps and what happens when it is incomplete or contradictory.
What belongs beside a percentage in my notes?
The denominator, selection rule, input conditions, outcome definition and software version. Those details decide which comparisons are meaningful.
Limits of this interpretation
- Selected benchmark cases and deliberately misleading inputs do not estimate errors during routine care.
- This source check covered selected indexed sections, not supplements or a reanalysis.
Sources & transparency
- Buckley, Diao, Srivastava et al. (2026): Multimodal foundation models exploit text to make medical image predictions
Publisher-indexed results and figure 2 caption, plus the primary paper’s PMC abstract and publication record, checked. Direct publisher retrieval failed; supplements, code and raw data were not checked. Only a short factual account is used; journal figures are not reproduced. · Accessed 28 Sep 2026
DOI: 10.1038/s41467-026-74207-5 - Bean, Payne, Parsons et al. (2026): Reliability of LLMs as medical assistants for the general public
Current publisher HTML: abstract, results, discussion and methods checked. Supplementary files and raw data were not independently assessed. The page links a 17 April 2026 correction; the separate correction text could not be retrieved. Our wording is original; this source is CC BY 4.0. · Accessed 28 Sep 2026
DOI: 10.1038/s41591-025-04074-y
Prepared and source-checked with AI on 28 September 2026 (UTC). Press-news Team is our collective publication byline, not a claim of medical credentials or human review. No human editorial or clinical review has taken place. This is an educational account of how AI was evaluated, not medical advice or a clinical endorsement. Access limitations are listed for each source; proposed follow-up tests are our commentary, not reported findings.
Source check: AI source check — selected primary research and reported numbers
Clinical review: Not applicable to this educational guide
Suggest a correction