A test can detect most cases and correctly classify most people without a condition, yet many positive results can still be false positives. The apparent puzzle disappears when you count the people in each group before turning the counts into percentages.
- Sensitivity starts with people who have the target condition; specificity starts with those who do not.
- Positive predictive value starts with positive results.
- A test’s performance must be assessed for its intended population and use.
The positive results in the 1% population
People among 585 positive results
Write down the four boxes
Start with a hypothetical test used in 10,000 people. Assume a reliable reference standard identifies 100 with the target condition and 9,900 without it. Set sensitivity to 90% and specificity to 95% for this exercise.
Among the 100 with the condition, 90 test positive and 10 negative. Among the 9,900 without it, 9,405 test negative and 495 positive. The four boxes are true positive, false negative, true negative and false positive, always defined relative to the reference standard.
These are deliberately invented counts. They describe no commercial test, disease prevalence or recommendation to seek testing. Their purpose is to make the denominators visible.
Now follow only the positive results
There are 90+495=585 positive results. Of those, 90 come from people with the condition. Positive predictive value is therefore 90/585, about 15.4%. The 90% sensitivity is not a 90% probability that a positive result is a true positive.
The FDA diagnostic guidance defines these measures and stresses the role of the population and reference standard. Our calculation simply applies the definitions to an original example.
Imagine a headline that replaces “90% sensitivity” with “90% accurate at telling you whether you have the disease.” The denominators have changed mid-sentence. A reader should ask which quantity the reported percentage actually describes.
Source 1 ↗Change the starting population, keep the assumed test
In a second hypothetical population of 10,000, let 1,000 have the condition. Keep sensitivity and specificity fixed at the same assumed values. We now have 900 true positives, 100 false negatives, 450 false positives and 8,550 true negatives.
Of 1,350 positive results, 900 are true positives: about 66.7%. No improvement in the assumed test was needed. Changing the starting frequency changed the composition of the positive results.
In real settings, sensitivity and specificity may also change with case mix, severity, thresholds and how testing is performed. Holding them constant here isolates one arithmetic relationship; it is not a guarantee that performance transfers unchanged between populations.
Check what counted as the truth
The reference standard is the method used to determine whether the target condition is present. If the comparator is itself imperfect, agreement with it is not automatically diagnostic accuracy. FDA guidance distinguishes agreement with a non-reference method from sensitivity and specificity against an appropriate standard.
For the fictional study, ask whether everyone received the reference assessment, whether its interpretation was independent of the new test and how uncertain results were handled. If only positive tests receive further investigation, the study may leave some false negatives undiscovered.
Do not silently remove failed or indeterminate tests from your summary. A test that cannot produce a usable result in some people has an operational limitation the reader needs to see.
Source 1 ↗Connect detection with the next step
An accuracy study answers how a test classifies people under specified conditions. A separate question is whether using the test in a care pathway improves outcomes, reduces burden or creates avoidable harm.
For our fictional test, map what happens after each result. A positive result might trigger another assessment; a negative result might delay one. The consequences depend on the condition, available follow-up and the costs of both kinds of error. Our example cannot choose that pathway.
When connecting papers, label one “classification performance” and another “effect of using the test.” The screening guide explains why finding more abnormalities is not by itself evidence that people live longer or feel better.
What would make the next validation useful?
A useful next study would recruit people resembling the intended users, define the threshold in advance, report all four boxes and explain inconclusive results. Independent evaluation in a new setting can test whether initial performance travels.
For our arithmetic example, the uncertainty around sensitivity and specificity is intentionally omitted because we have assigned exact teaching values rather than estimated them. A real study needs uncertainty intervals and enough observations of both condition-positive and condition-negative people.
Your reading conclusion should name the test’s job, population, comparator and uncertainty. “Promising classification performance in this study population” can be defensible while a blanket claim that the test diagnoses everyone reliably remains unsupported.
Two populations, the same assumed sensitivity and specificity
| Count or measure | Condition in 1% | Condition in 10% |
|---|---|---|
| People tested | 10,000 | 10,000 |
| True positives | 90 | 900 |
| False negatives | 10 | 100 |
| False positives | 495 | 450 |
| True negatives | 9,405 | 8,550 |
| All positive results | 585 | 1,350 |
| Positive predictive value | 90/585 ≈ 15.4% | 900/1,350 ≈ 66.7% |
Original hypothetical counts; sensitivity 90%, specificity 95%, complete verification by an assumed reliable reference standard. Not observed clinical performance.
Your questions, answered
Is sensitivity the probability of disease after a positive result?
No. Sensitivity is the proportion testing positive among those with the condition. Positive predictive value is the proportion with the condition among those testing positive. The two reverse the denominator.
Why are there so many false positives in the first example?
The group without the condition is much larger. Even a 5% false-positive proportion among 9,900 people produces 495 results, compared with 90 true positives among the 100 with the condition.
Can a negative result rule out every possibility?
No. In our first fictional population, ten people with the condition test negative. What a negative result means depends on the test, context and prior information; these examples do not guide an individual diagnosis.
Do sensitivity and specificity always stay constant?
No. They can vary with the population, disease spectrum, threshold and measurement process. We hold them fixed only to demonstrate how starting frequency changes predictive value.
What if the study uses another test as the comparator?
Find out whether it is a defensible reference standard. Agreement with an imperfect non-reference test is not automatically sensitivity or specificity for the underlying condition.
What evidence comes after a good accuracy study?
Evidence about how the test performs in intended settings and what happens when it is used in a care pathway. Better classification and better health outcomes are connected questions, but they are not interchangeable.
Limits of this interpretation
- This is a selected educational explanation, not a systematic review, validated appraisal instrument or personal care recommendation.
- Numerical examples are hypothetical. Their deliberately simplified assumptions must not be transferred to a real study without checking its methods.
- An AI source check can miss errors; source access and the absence of independent human review are stated explicitly.
Sources & transparency
- FDA: Statistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests (2007)
Public full HTML guidance; definitions, predictive values, intended-use population and reference-standard cautions checked by AI. All counts in our worked examples are invented. · Accessed 27 Sep 2026
Prepared and source-checked with AI on 27 September 2026. Press-news Team is the publication’s collective byline, not a claim of medical credentials or human review. No human editorial or clinical review has taken place. Source access is described below each reference. Worked examples are invented for education and do not report a clinical trial or predict an individual outcome.
Source check: AI source check — selected methods references and worked examples
Clinical review: No human editorial or clinical review
Suggest a correction