Sensitivity and specificity look like they can be computed from any 2×2 table comparing a new test to an existing one. They cannot, at least not correctly. Both are defined against truth: the true disease status of each subject. If the thing you compared your new test against is not a true reference standard but only another imperfect test, you have not measured sensitivity at all. You have measured agreement. The correct terms are positive and negative percent agreement.
Sensitivity is the proportion of truly diseased subjects the test flags. To compute it you need to know, independently, who is truly diseased. That means a reference standard: biopsy, culture, long-term follow-up, an established definitive method. Often no such standard is available. The reference method may be too invasive to use on every subject, or it may be imperfect itself. So a new assay gets compared against the best existing test instead.
The moment your comparator is another fallible test rather than the truth, the denominators change meaning. Evaluating a qualitative assay end to end covers where this sits in the wider study. “Of everyone the comparator called positive, how many did the new test also call positive?” is not sensitivity. The comparator’s positives are not the truly diseased. Calling it sensitivity credits the new test with detecting disease when all you have shown is that it agrees with another test that is itself sometimes wrong.
The correct statistics for this situation are positive percent agreement (PPA) and negative percent agreement (NPA).
PPA = (both tests positive) / (all comparator positives)
NPA = (both tests negative) / (all comparator negatives)
Arithmetically they look just like sensitivity and specificity. Same cell counts, same ratios. But the name is a deliberate signal. The denominator is a comparator method, not a reference standard. No claim about detecting true disease is being made.
The distinction is not pedantry. The US FDA’s guidance on reporting diagnostic test results is explicit. When the comparator is not a reference standard, results should be reported as percent agreement rather than sensitivity and specificity. That way readers are not misled into reading an agreement study as a proof of accuracy. Performance studies for an IVD 510(k) sets out the rest of what that guidance recommends.
Two tests can agree beautifully and both be wrong in the same way. High PPA and NPA tell you the new test behaves like the comparator, which is useful if the comparator is trusted and you are seeking equivalence. High agreement says nothing about whether either test is actually right. Reporting that as “98% sensitivity” makes a claim about disease detection the study never supported. In a regulatory submission it is the kind of overstatement that gets the submission rejected.
The wording also changes how a reader should act. Sensitivity permits statements about ruling disease out. Percent agreement permits statements about interchangeability with the comparator. Those are different clinical claims, and the words carry the difference.
| Your comparator is… | Report… |
|---|---|
| A true reference standard (biopsy, culture, definitive method, confirmed status) | Sensitivity and specificity |
| Another imperfect test or predicate method (no independent truth) | Positive and negative percent agreement |
The test is simple. Ask whether your comparator establishes the truth or only represents current practice. If you could not rely on the comparator being right in every case, you are measuring agreement. PPA/NPA is what you report.
Each agreement measure is a proportion. Each needs a confidence interval. The Wilson score interval is the sound default, especially with the small positive counts common in early evaluations. When you are comparing the agreement of two methods, or a difference in agreement, the Newcombe interval handles the difference between two proportions properly. As always, a bare percentage without its interval hides how much the sample could be misleading you.
Download the CLSI EP12-A2 method comparison example workbook (.xlsx) — two qualitative methods for the same analyte compared, with positive and negative percent agreement and Wilson confidence intervals, ready to open in the Analyse-it trial.
The example workbook is downloading.
It opens in Excel on its own — the data and the finished results are both in it. Analyse-it is what lets you change the analysis and re-run it, try the same study on your own data, or work through it to see how the software handles it.
Every feature from all five editions for 15 days.
Reporting sensitivity when you mean PPA. If the denominator is a comparator method rather than confirmed truth, the number is percent agreement. Naming it sensitivity claims something the study cannot support.
Reading high agreement as proof of accuracy. Two tests can agree and both be wrong. Agreement shows equivalence to the comparator, not correctness.
Resolving discordants selectively. Re-testing only the cases where the two disagree, with a third method, and then including those results, biases the estimate. If you resolve discordants, do it by a pre-specified rule and report the unresolved figures too.
Dropping the confidence interval. Early agreement studies often have few positives. The interval on PPA can be wide. Report it.
Analyse-it evaluates qualitative tests per EP12-A2, and the wording follows the comparator, inside Excel:
Every feature from all five editions for 15 days. Qualitative test evaluation is in the Method Validation and Ultimate editions, from US$ 475 a year. Validated against NIST and CLSI reference datasets. See the sensitivity and specificity guide for the reference-standard case; percent agreement counts raw concordance, while Cohen’s kappa corrects it for chance.