A study can measure many outcomes, at many visits, in many groups. Somewhere in that collection, an eye-catching p-value may appear. To understand it, you need the map of the questions that were asked, not only the number that was selected for the abstract.

THE SHORT READ
  • A p-value is not the probability that a treatment works.
  • The number and organisation of tests affect the chance of false-positive claims.
  • Exploratory findings can be valuable when their status stays visible.
THE NUMBERS, IN CONTEXT

More independent chances to cross a threshold

Probability (%) of at least one false positive

1 test5
10 tests40.1
20 tests64.2
50 tests92.3
0100
Original hypothetical calculation: all null hypotheses true, independent valid tests, 5% per-test threshold. Full 0–100% scale; not a measured rate in published research.

Put the p-value back inside its question

A p-value describes how unusual the observed test statistic, or something more extreme, would be under a specified null model and its assumptions. It does not tell you the probability that the hypothesis is true, the size of a benefit or whether the result matters to patients.

The ASA statement warns against letting a threshold carry a scientific conclusion by itself. For a reader, the practical response is to copy four things together: the effect estimate, uncertainty interval, tested comparison and analysis status.

Imagine “p=0.03” printed without any of those companions. You cannot yet tell whether the study concerns a meaningful outcome or a minor measurement chosen from a long list.

Source 1 ↗

An invented experiment with twenty chances

Assume twenty independent tests, all null hypotheses true, with each test having an exact 5% chance of a false positive. The chance of no false positives is 0.95 multiplied by itself twenty times. The chance of at least one is 1−0.95²⁰, approximately 64.2%.

This is our own probability example. It is not an estimate of how often published medical papers are wrong. Real endpoints are often correlated; the simple independence calculation then does not apply directly.

The example shows why “we found one p-value below 0.05” can mean something very different after one planned test and after twenty opportunities. It does not imply that the one selected finding has a 64.2% probability of being false.

Find the planned order of questions

FDA’s multiplicity guidance describes ways to organise endpoint families and control false-positive conclusions. In a paper, look for the protocol or statistical analysis plan explaining which tests were confirmatory and how multiple comparisons were handled.

Our fictional trial names walking distance at week 12 as primary, with five secondary outcomes tested in a planned sequence. The primary result does not meet the specified criterion, but a later fatigue comparison has an unadjusted p-value of 0.02. Whether that later result supports a confirmatory claim depends on the actual testing strategy; its small number alone does not restart a stopped sequence.

Record “nominal” or “unadjusted” when the authors use those terms. Do not silently drop them from a summary.

Source 2 ↗

Exploration is useful when the label travels with it

Suppose a team notices that fatigue improves only in participants who started with poor sleep. That could suggest a worthwhile new study. The question is whether this pattern was predicted and tested as planned, or discovered while examining the data.

Write a future hypothesis in a separate sentence: “A subsequent trial could test whether baseline sleep changes the effect on fatigue.” That sentence keeps the idea alive without converting it into a treatment rule.

An exploratory analysis can still contain careful methods and an important observation. The reporting problem begins when discovery is presented as though it was an independent confirmation of a prediction made beforehand.

A correction does not repair every weakness

In the twenty-test exercise, a simple Bonferroni rule would use 0.05/20=0.0025 for each test to control the chance of at least one false positive at no more than 5%. Unlike our exact probability formula, this control does not require independent tests, provided the individual tests are valid.

That is an educational illustration, not a recommendation to apply the same rule mechanically to every study. Planned hierarchies and other procedures answer different design needs.

A multiplicity adjustment also cannot rescue biased outcome collection or a clinically unhelpful endpoint. After checking the testing plan, return to the outcome and effect size. Statistical organisation is one part of the argument.

Source 2 ↗

What a useful replication would look like

For the fictional fatigue finding, a useful follow-up would specify the target population, outcome, assessment time and subgroup prediction before results are examined. It would explain the main comparison and publish the result even if the pattern fails to recur.

Do not treat a second analysis of the same participants as an independent replication. It may test robustness, but shares the original observations. A genuinely new test would add information of a different kind.

Before sharing a headline, try this revision: “The study generated a fatigue hypothesis among several exploratory analyses; a planned independent test is still needed.” The useful news is the question that emerged, with the evidence stage made explicit.

CONNECT THE EVIDENCE

Our independent-tests calculation

Number of testsChance of at least one false positiveAssumptions
15.0%All null hypotheses true; independent valid tests at 5%
1040.1%The same assumptions
2064.2%The same assumptions
5092.3%The same assumptions

Computed as 100 × (1 − 0.95^m), rounded to one decimal. Not a probability that any particular study conclusion is false.

READER QUESTIONS

Your questions, answered

Does p=0.03 mean a 97% chance that the treatment works?

No. The p-value is calculated under a specified null model; it is not a probability assigned to the treatment hypothesis. It also does not measure the magnitude or importance of the effect.

Is 0.049 fundamentally different from 0.051?

A prespecified decision rule may place them on different sides of a threshold. Their scientific interpretation should still consider the nearby estimates, uncertainty, design and context rather than treating them as opposite worlds.

Does the 64.2% example describe real clinical research?

No. It assumes twenty independent valid tests and twenty true null hypotheses. It demonstrates a mathematical possibility under explicit conditions, not the false-positive rate of medical literature.

Are secondary outcomes unimportant?

No. They can be highly important. Check whether their analyses support confirmatory claims under the planned multiplicity strategy, or whether the results are exploratory or descriptive.

Should every exploratory finding be ignored?

No. It can guide a better question and a new study. Preserve its exploratory status, describe the search that produced it and avoid presenting it as a confirmed effect.

What should I ask when the testing plan is unavailable?

Ask which comparison was primary, how many outcomes and time points were examined, and whether the reported p-values were adjusted. Record unavailable information explicitly rather than assuming a plan existed.

LIMITATIONS

Limits of this interpretation

  • This is a selected educational explanation, not a systematic review, validated appraisal instrument or personal care recommendation.
  • Numerical examples are hypothetical. Their deliberately simplified assumptions must not be transferred to a real study without checking its methods.
  • An AI source check can miss errors; source access and the absence of independent human review are stated explicitly.
SOURCE NOTES

Sources & transparency

  1. American Statistical Association: Statement on statistical significance and p-values (2016)

    Public three-page ASA release and its six principles checked by AI; not a claim to have reviewed every accompanying commentary. · Accessed 27 Sep 2026

  2. FDA: Multiple Endpoints in Clinical Trials (2022)

    Public final guidance PDF, selected passages on multiplicity, endpoint families and ordering of tests checked by AI. The probability calculations in this article are original teaching examples. · Accessed 27 Sep 2026

Prepared and source-checked with AI on 27 September 2026. Press-news Team is the publication’s collective byline, not a claim of medical credentials or human review. No human editorial or clinical review has taken place. Source access is described below each reference. Worked examples are invented for education and do not report a clinical trial or predict an individual outcome.

Source check: AI source check — selected methods references and worked examples

Clinical review: No human editorial or clinical review

Suggest a correction