Explanations are easy to mistake for verification. The more interesting educational question is whether a learner can recognize when an explanation deserves trust. Accuracy and confidence need separate measurements.
- Paper
- npj Digital Medicine · 17 March 2026
- Design
- 111 students; three randomized groups
Original publication: 17 Mar 2026 · The date above refers to this brief.
Correct answers by assigned condition
Question responses answered correctly (%)
The study in brief
Teng and colleagues compared no explanation, correct explanations and deliberately misleading explanations across 2,775 question responses. Accuracy was 21.0%, 22.6% and 9.2%, respectively. Compared with control, the misleading condition had an odds ratio of 0.36 (95% CI 0.27–0.47); correct explanations showed no significant accuracy gain. Both explanation groups reported higher overall confidence. The explanations were prepared and expert-checked for the experiment, not an unfiltered sample of everyday chatbot replies.
Source 1 ↗Confidence and calibration are different questions
Consider an invented quiz with ten answers. A learner who is certain about every response has high confidence. A learner who expresses greater confidence on correct answers than on incorrect ones has useful discrimination. Neither description alone tells you whether an “80% sure” statement is correct about 80% of the time.
When reading a calibration claim, look for its operational definition. Was confidence a probability, a category or a rating? Did the analysis compare average confidence, separate right from wrong answers, or test numerical agreement between confidence and observed accuracy? Those are related but different questions.
A controlled error is not an everyday error rate
An experiment can deliberately introduce misinformation to test susceptibility. That design does not tell us how frequently users encounter comparable misinformation in ordinary use. To estimate everyday impact, we would also need to know how often errors appear, whether people notice them, and what follows after they do.
Our suggested reading note has two boxes: the effect of exposure under experimental conditions, and the frequency of that exposure outside the experiment. Leaving the second box empty is more informative than filling it with a guess.
A related result, in a different population
The separate Bean study tested adults using chatbots to reason through written health scenarios. Their next-step choices were not significantly better than control, despite strong model-only performance. It helps frame a common evaluation question: what changes in the person’s answer? It does not establish that students and the public respond identically or that the same mechanism explains both findings.
Source 2 ↗What a useful next experiment could measure
Our proposal is to compare learning strategies, not simply access to an explanation. For example, learners could record an answer and uncertainty first, inspect evidence second, and explain any revision third. A comparison group would complete an otherwise equivalent task.
A preregistered follow-up could test delayed recall, recognition of deliberately planted errors and transfer to unfamiliar questions. It should record whether assistance improves learning when the tool is removed. These are hypotheses for further study, not proven remedies or a curriculum recommendation.
Three claims to keep separate
| Claim | Evidence you would want |
|---|---|
| The learner feels more certain | A clearly defined confidence measure |
| The learner answers more accurately | A scored outcome and a fair comparison |
| The learner knows when to doubt | Confidence examined alongside correctness |
| Learning lasts after assistance ends | A later assessment without the tool |
Original reading framework, not extra findings from the trial.
Your questions, answered
Does a convincing explanation independently verify an answer?
No. Rephrasing a claim, extending it or making it sound orderly is not an independent check against evidence.
Why distinguish students from practicing clinicians?
Expertise changes the task a reader brings to an explanation. Generalizing between groups needs evidence rather than an assumption.
What should I record when reading an AI education study?
Who learned, what assistance they received, what was measured immediately, and whether anything was measured later without assistance.
Limits of this interpretation
- A short educational task does not establish patient outcomes, durable learning or everyday misinformation frequency.
- This source check did not independently examine supplements or registration details.
Sources & transparency
- Teng, Tan, Cao et al. (2026): Impact of AI misinformation on diagnostic accuracy and confidence calibration in novice medical students
Publisher-indexed abstract, results and study-design sections checked. Direct full-page retrieval failed; supplements, participant data and registration record were not independently assessed. Only a short factual account is used. · Accessed 28 Sep 2026
DOI: 10.1038/s41746-026-02547-z - Bean, Payne, Parsons et al. (2026): Reliability of LLMs as medical assistants for the general public
Current publisher HTML: abstract, results, discussion and methods checked. Supplementary files and raw data were not independently assessed. The page links a 17 April 2026 correction; the separate correction text could not be retrieved. Our wording is original; this source is CC BY 4.0. · Accessed 28 Sep 2026
DOI: 10.1038/s41591-025-04074-y
Prepared and source-checked with AI on 28 September 2026 (UTC). Press-news Team is our collective publication byline, not a claim of medical credentials or human review. No human editorial or clinical review has taken place. This is an educational account of how AI was evaluated, not medical advice or a clinical endorsement. Access limitations are listed for each source; proposed follow-up tests are our commentary, not reported findings.
Source check: AI source check — selected primary research and reported numbers
Clinical review: Not applicable to this educational guide
Suggest a correction