Several AI agents can reach the same answer. A benchmark can tell us whether that answer matches a reference. Neither observation, by itself, tells us whether an evolving evidence service improves care.
- Name the reference and the unit before calling a system accurate.
- Agent agreement is different from independent confirmation.
- Updating evidence and matching an existing guideline can pull in different directions.
- Publication
- Communications Medicine · 7 October 2026
- Version
- Accepted early version; subject to edits
- Coverage here
- Benchmark interpretation only
Original publication: 7 Oct 2026 · The date above refers to this brief.
The reported benchmark
Wang and colleagues evaluated MARG against seven 2025 liver-disease guidelines across 137 scenarios, reporting 77.4% overall concordance. The publisher describes an accepted early version. Our coverage uses the available HTML summary and AI-use statement; the full methods and scenario records were not assessed.
Source 1 ↗Concordance needs a named reference
Agreement is interpretable only when the reference is visible. Does the score compare a binary recommendation, its supporting evidence, or a complete reasoning chain? Does each scenario count once, or are outputs weighted?
Those details are questions for the full evaluation record, not answers we can reconstruct from a headline percentage. We therefore preserve the author-reported value and do not convert it into a count of correct decisions. The denominator in an abstract does not necessarily specify the scoring unit.
Consensus can repeat a shared mistake
Imagine several agents that draw from the same repository and use similar prompts. Their agreement could reflect a well-supported conclusion, but it could also repeat the same missing source or interpretation error. This is a hypothetical failure mode, not an observed failure assigned to this system.
Our proposed check would compare consensus with an independently adjudicated reference and retain disagreement cases for inspection. Count who actually supplied independent evidence. The number of software roles should not become the number of independent expert confirmations.
Updating and matching are different goals
A service designed to find new evidence may sometimes diverge from an older reference. A service designed to reproduce that reference may score well while missing an important update. Evaluating both goals with a single agreement percentage hides that tension.
Our suggested design freezes a reference date, labels later publications separately and distinguishes faithful reproduction from justified change. An audit trail should show which source changed, which recommendation it affects and who reviewed the revision. These are proposed safeguards and measurements, not a description of controls verified in the available summary.
A concrete next benchmark
We would construct held-out cases with missing, conflicting, superseded and retracted evidence, plus stable cases where no update is warranted. Before running the system, define whether the expected action is to retain, revise or defer an answer.
Measure source retrieval, evidence interpretation and escalation separately. Record inappropriate changes as well as missed changes. This future test would make maintenance quality measurable without treating a polished answer as a successful evidence update. It would still be a workflow evaluation, rather than proof of better patient outcomes.
Connect the model with the service
Read this alongside our MedGemma and human-user chatbot explainers. They help separate a component’s performance from the behavior of an entire service. Our cross-paper question is where an error can enter and who has a practical way to catch it.
The hopeful direction is an evidence tool whose output can be traced and challenged. Demonstrating that direction requires accessible evaluation records and independent tests. This article makes no liver-disease treatment recommendations and does not assess whether the generated recommendations are clinically appropriate.
Agreement with what, counted how?
Accepted early-version HTML only. Numbers below are author-reported scope and concordance, not independently recalculated clinical accuracy.
01Does the score measure better care?
- What was observed
- The benchmark uses guideline references.
- Where the conclusion stops
- No patient-benefit estimate follows.
Source 1 · Abstract → Methods and Results
02Are seven agents seven independent reviewers?
- What was observed
- The disclosure identifies seven AI roles.
- Where the conclusion stops
- Role count does not establish independence.
Source 1 · Statement on using artificial intelligence
Numbers you can inspect
| Measure | Value & unit | Origin & method |
|---|---|---|
| Reference guidelines | 7 guidelines | Reported 2025 reference set. Source 1 · Abstract → Methods |
| Evaluation coverage | 137 scenarios | Reported Not treated patients; scoring records unavailable. Source 1 · Abstract → Methods |
| Overall concordance | 77.4 percent | Reported Author-reported; not converted into a correct-case count. Source 1 · Abstract → Results |
Compare the actual experiments
These studies answer different questions. Read the unit and endpoint before comparing results.
| Study | Unit & setting | Readout | Interpretation boundary |
|---|---|---|---|
| Wang et al. 2026 Source 1 · Abstract → Methods and Results | Scenario outputs | Reference concordance | Not clinical benefit. |
The export includes claims, available numbers, methods and source locations. It contains our reading notes and published summaries; it is not raw participant data or an independent reanalysis.
Evidence update · 7 Oct 2026
First publication. Reported numbers retain their units and source locations. Calculations are labelled. Interpretation and proposed experiments are ours; no participant data or experimental records were reanalysed. AI source check; no human editorial or clinical review.
Four evaluations to keep separate
| Evaluation | Question | Additional evidence needed |
|---|---|---|
| Guideline concordance | Does output match a dated reference? | Scoring rules and scenario-level adjudication |
| Agent consensus | Do software roles converge? | Independence and shared-error tests |
| Evidence maintenance | Was an update justified? | Versioned sources and reviewed change records |
| Real-world benefit | Does the intended service improve outcomes? | An appropriate prospective comparison |
Our proposed evaluation framework. The rows are not four benefits demonstrated in the paper.
Your questions, answered
Is 77.4% a patient-success rate?
No. It is reported guideline concordance.
Can consensus replace an independent reviewer?
Agreement among software roles does not demonstrate independent scrutiny. The reviewer’s task and evidence would need to be specified.
Why not calculate the number of correct scenarios?
The scoring unit, weighting and rounding need full-method verification. Multiplying a summary percentage would create unwarranted precision.
Why record the publication version?
An accepted early version may be edited. The source version makes a later correction or update traceable.
Limits of this interpretation
- Only the accepted-version HTML summary and disclosure were assessed.
- Scenario-level correctness and weighting were not independently audited.
- Guideline concordance is not a measured patient outcome.
- Consensus does not by itself establish independent validation.
Sources & transparency
- Wang, Wen, Xue et al. (2026): Multi-agent AI for liver disease guideline development and maintenance: a proof-of-concept study
Accepted early-version publisher HTML abstract, plain-language summary, version notice and AI-use statement checked. Primary PDF retrieval failed; full methods, scenario-level records, guideline grading and supplements were not inspected. Coverage is limited to evaluation design and reported overall concordance, not clinical recommendations or treatment appraisal. · Accessed 7 Oct 2026
DOI: 10.1038/s43856-026-01961-4
Prepared and source-checked with AI. Press-news Team is the collective publication byline, not a medical reviewer. No human editorial or clinical review has taken place. This educational article explains research methods and basic research; it does not provide individual diagnosis or treatment recommendations. We did not conduct these experiments or reanalyse participant data. Findings, interpretation and proposed follow-up tests are distinguished. Source access is recorded below. Photographs are illustrative.
Source check: AI source check — named primary passages, measurement units and interpretation boundaries
Clinical review: Not applicable to this educational guide
Suggest a correction