Explanations are easy to mistake for verification. The more interesting educational question is whether a learner can recognize when an explanation deserves trust. Accuracy and confidence need separate measurements.

THE GUIDE AT A GLANCETeng et al. · Educational experiment
Paper
npj Digital Medicine · 17 March 2026
Design
111 students; three randomized groups

Original publication: 17 Mar 2026 · The date above refers to this brief.

THE NUMBERS, IN CONTEXT

Correct answers by assigned condition

Question responses answered correctly (%)

No explanation21
Correct explanation22.6
Misleading explanation9.2
0100
Raw percentages; repeated responses within students/questions. Point estimates only, not confidence intervals or patient outcomes.

The study in brief

Teng and colleagues compared no explanation, correct explanations and deliberately misleading explanations across 2,775 question responses. Accuracy was 21.0%, 22.6% and 9.2%, respectively. Compared with control, the misleading condition had an odds ratio of 0.36 (95% CI 0.27–0.47); correct explanations showed no significant accuracy gain. Both explanation groups reported higher overall confidence. The explanations were prepared and expert-checked for the experiment, not an unfiltered sample of everyday chatbot replies.

Source 1 ↗

Confidence and calibration are different questions

Consider an invented quiz with ten answers. A learner who is certain about every response has high confidence. A learner who expresses greater confidence on correct answers than on incorrect ones has useful discrimination. Neither description alone tells you whether an “80% sure” statement is correct about 80% of the time.

When reading a calibration claim, look for its operational definition. Was confidence a probability, a category or a rating? Did the analysis compare average confidence, separate right from wrong answers, or test numerical agreement between confidence and observed accuracy? Those are related but different questions.

A controlled error is not an everyday error rate

An experiment can deliberately introduce misinformation to test susceptibility. That design does not tell us how frequently users encounter comparable misinformation in ordinary use. To estimate everyday impact, we would also need to know how often errors appear, whether people notice them, and what follows after they do.

Our suggested reading note has two boxes: the effect of exposure under experimental conditions, and the frequency of that exposure outside the experiment. Leaving the second box empty is more informative than filling it with a guess.

A related result, in a different population

The separate Bean study tested adults using chatbots to reason through written health scenarios. Their next-step choices were not significantly better than control, despite strong model-only performance. It helps frame a common evaluation question: what changes in the person’s answer? It does not establish that students and the public respond identically or that the same mechanism explains both findings.

Source 2 ↗

What a useful next experiment could measure

Our proposal is to compare learning strategies, not simply access to an explanation. For example, learners could record an answer and uncertainty first, inspect evidence second, and explain any revision third. A comparison group would complete an otherwise equivalent task.

A preregistered follow-up could test delayed recall, recognition of deliberately planted errors and transfer to unfamiliar questions. It should record whether assistance improves learning when the tool is removed. These are hypotheses for further study, not proven remedies or a curriculum recommendation.

CONNECT THE EVIDENCE

Three claims to keep separate

ClaimEvidence you would want
The learner feels more certainA clearly defined confidence measure
The learner answers more accuratelyA scored outcome and a fair comparison
The learner knows when to doubtConfidence examined alongside correctness
Learning lasts after assistance endsA later assessment without the tool

Original reading framework, not extra findings from the trial.

READER QUESTIONS

Your questions, answered

Does a convincing explanation independently verify an answer?

No. Rephrasing a claim, extending it or making it sound orderly is not an independent check against evidence.

Why distinguish students from practicing clinicians?

Expertise changes the task a reader brings to an explanation. Generalizing between groups needs evidence rather than an assumption.

What should I record when reading an AI education study?

Who learned, what assistance they received, what was measured immediately, and whether anything was measured later without assistance.

LIMITATIONS

Limits of this interpretation

  • A short educational task does not establish patient outcomes, durable learning or everyday misinformation frequency.
  • This source check did not independently examine supplements or registration details.
SOURCE NOTES

Sources & transparency

  1. Teng, Tan, Cao et al. (2026): Impact of AI misinformation on diagnostic accuracy and confidence calibration in novice medical students

    Publisher-indexed abstract, results and study-design sections checked. Direct full-page retrieval failed; supplements, participant data and registration record were not independently assessed. Only a short factual account is used. · Accessed 28 Sep 2026

    DOI: 10.1038/s41746-026-02547-z
  2. Bean, Payne, Parsons et al. (2026): Reliability of LLMs as medical assistants for the general public

    Current publisher HTML: abstract, results, discussion and methods checked. Supplementary files and raw data were not independently assessed. The page links a 17 April 2026 correction; the separate correction text could not be retrieved. Our wording is original; this source is CC BY 4.0. · Accessed 28 Sep 2026

    DOI: 10.1038/s41591-025-04074-y

Prepared and source-checked with AI on 28 September 2026 (UTC). Press-news Team is our collective publication byline, not a claim of medical credentials or human review. No human editorial or clinical review has taken place. This is an educational account of how AI was evaluated, not medical advice or a clinical endorsement. Access limitations are listed for each source; proposed follow-up tests are our commentary, not reported findings.

Source check: AI source check — selected primary research and reported numbers

Clinical review: Not applicable to this educational guide

Suggest a correction