Two treatments can look close while an important difference remains unresolved. A non-inferiority result needs a justified boundary, a direction of comparison and an interval that stays on the right side of that boundary.
- Find the pre-specified margin and the reason it was chosen.
- For this example, positive differences favour standard care; an upper interval bound below +4 meets the numerical non-inferiority rule.
- An interval can exclude the unacceptable margin while still favouring standard care.
- Passing the numerical rule alone does not establish that the design and assumptions are credible.
- Outcome in the example
- An undesirable event within one year
- Effect measure
- New-strategy risk minus standard-care risk
- Illustrative boundary
- +4 percentage points; not a clinical recommendation
- Uncertainty
- Invented two-sided 95% confidence intervals
The question behind the design
A non-inferiority trial asks whether a new treatment avoids a loss of efficacy larger than a pre-specified margin against an active comparator. The margin needs clinical justification and evidence about the comparator’s established effect. Close results alone are insufficient: the comparator must be expected to work in the new setting. This dependence on historical evidence and trial quality is central to the design.
Source 1 ↗Set up the ruler before reading the result
For the following exercise, imagine an undesirable event within one year. We subtract the standard-care risk from the new-strategy risk. A positive difference means more events with the new strategy. A negative difference means fewer. Zero means equal estimated risks.
We invent a boundary of +4 percentage points and use two-sided 95% confidence intervals. In this exercise, the upper bound must be below +4 to meet the numerical non-inferiority rule. The boundary is chosen solely to make the arithmetic readable; it would not be an acceptable justification for any real trial. These intervals are illustrations, not estimates calculated from participant data.
If the standard-care risk were 10%, an extra 4 percentage points would mean 14%, not 10.4%. That is 40 extra events per 1,000 people. If the starting risk were 2%, the same absolute difference would mean 6%. Those contrasting translations show why a naked “4% margin” is unhelpful without a scale and clinical context.
Four results, four different readings
Scenario A estimates a difference of +1 point, with an interval from −1 to +3. Its upper bound stays below +4. It meets our numerical rule while leaving both a small benefit and a small disadvantage compatible with the interval. It does not establish identical effects.
Scenario B has exactly the same +1 point estimate but an interval from −1 to +6. This fails to exclude our boundary. The result is inconclusive for non-inferiority. A headline based only on the central estimate would miss the difference between A and B.
Scenario C estimates −2 points, with an interval from −3 to −1. It stays below +4 and below zero. In a suitable pre-specified testing plan, that pattern supports superiority as well as non-inferiority.
Scenario D estimates +2 points, with an interval from +1 to +3. Every value in the interval favours standard care, yet the interval remains below the +4 boundary. The numerical criteria for a difference and non-inferiority can therefore both be satisfied. “Non-inferior” cannot safely be translated as “equally effective”.
Source 2 ↗Why the interval is only part of the story
The trial must be capable of detecting a meaningful difference if one exists: this is assay sensitivity. Expectations based on older comparator trials must remain plausible in the new trial, often called the constancy assumption. Poor adherence, treatment crossover or other conduct problems can make the groups look misleadingly similar. Intention-to-treat analysis is not automatically conservative for a non-inferiority question.
Source 1 ↗Read the analysis population, then the other outcomes
Look for results under the planned intention-to-treat and per-protocol approaches, with exclusions explained. Neither a convenient subset nor agreement between analyses removes every possible bias. The report should connect its conclusion to its original hypotheses and show how sensitive the result is to analytical choices.
A claim about one efficacy endpoint also leaves separate questions about adverse effects, burden and cost. If a new treatment is described as easier to use, look for evidence for that advantage rather than treating the description as established by the non-inferiority test.
Source 2 ↗Turn a result into a precise paragraph
For Scenario A, our editorial exercise would read: “The invented estimate was one additional event per 100 people, with an interval ranging from one fewer to three more. That interval excludes the exercise’s boundary of four additional events per 100. Whether such a boundary would be acceptable in practice has not been established.”
For Scenario B, only the second sentence changes substantially: “The interval extends to six additional events per 100, so it does not exclude the exercise’s boundary.” It would be inaccurate to write that B proved the new strategy inferior. It would be equally inaccurate to label B reassuring because its point estimate was close to zero.
For Scenario D, retain both pieces of information: the interval favours standard care, and the disadvantage falls inside the chosen boundary. Readers need the size of that possible trade-off, not just the name of the statistical conclusion.
What would make the evidence more useful?
For our hypothetical report, the next useful addition would be a documented rationale for the boundary. After that, we would want the underlying event counts, follow-up completeness, analyses of missing outcomes and evidence for any claimed convenience or safety advantage. Until those exist, the exercise remains a lesson about intervals.
When connecting two real trials, copy the outcome, effect scale and margin from each report before comparing their conclusions. A study testing an absolute risk difference and one testing a risk ratio cannot be compared by lining up the margin numbers alone. The question to carry forward is specific: which remaining uncertainty would a new study need to reduce?
Read the interval against zero and the margin
| Invented scenario | Difference (percentage points) | Invented 95% interval | Numerical reading in this exercise |
|---|---|---|---|
| A | +1 | −1 to +3 | Meets the +4 boundary rule; superiority not shown |
| B | +1 | −1 to +6 | Non-inferiority not established; inconclusive |
| C | −2 | −3 to −1 | Meets the rule; superiority also supported if properly tested |
| D | +2 | +1 to +3 | Meets the rule while the interval favours standard care |
Original hypothetical scenarios. New minus standard risk of an undesirable one-year outcome; positive is worse. Upper bound below +4 is the numerical rule used here. Neither the intervals nor the margin came from a real trial. Design assumptions still matter.
Your questions, answered
Does a non-significant difference prove non-inferiority?
No. Scenario B includes zero but also extends beyond the unacceptable boundary. Failure to establish a difference is not enough to establish non-inferiority.
Does failing non-inferiority prove the new treatment is worse?
Not necessarily. Scenario B remains compatible with a small benefit and a substantial disadvantage. The interval leaves the non-inferiority question unresolved.
Can a treatment be non-inferior and statistically worse?
Yes in a numerical comparison such as Scenario D. Its interval excludes zero in favour of standard care but stays below the chosen margin. Clinical acceptability is another question.
Is non-inferiority the same as equivalence?
No. Non-inferiority addresses one unacceptable direction. Equivalence uses boundaries in both directions. Neither term means that the effects are exactly identical.
Can I choose a margin after seeing the interval?
That would change the exercise into moving the goalposts. For example, replacing +4 with +7 would make Scenario B pass without adding any new evidence. A real margin needs advance justification.
Why not use these four examples to choose a treatment?
They contain no actual treatment, population or observed data. Their purpose is to help you read a methods section and recognise which question an interval can answer.
Limits of this interpretation
- The +4 percentage-point margin is a teaching device, not an acceptable threshold for any disease or treatment.
- No sample sizes, participant records or calculated confidence intervals are supplied; the four scenarios illustrate logical relationships.
- This guide does not establish equivalence, effectiveness, safety or interchangeability for any product.
Sources & transparency
- FDA: Non-Inferiority Clinical Trials to Establish Effectiveness (2016)
Final guidance PDF accessed. Sections on margin selection, assay sensitivity, constancy and study quality were checked. No treatment-specific margin is recommended by this guide. · Accessed 27 Sep 2026
- Piaggio and colleagues: CONSORT extension for noninferiority and equivalence trials (JAMA, 2012)
Accessible publisher full-text article, especially statistical methods and interpretation. The original reporting extension informs the explanation of intervals, analysis populations and pre-specified hypotheses. The scenarios below are newly invented. · Accessed 27 Sep 2026
Written and source-checked by AI using the accessible documents and sections identified below. No human editorial or clinical review has been completed. All worked examples are hypothetical and were created for this guide; they are not trial findings or treatment advice.
Source check: AI source check — 27 September 2026
Clinical review: Not applicable to this educational guide
Suggest a correctionA quick reference
Look up a term, work through an example and check a common pitfall.
