Positive and negative percent agreement vs sensitivity and specificity When there is no gold standard to compare against, sensitivity and specificity are the wrong words — and using them anyway overstates what your study showed. What positive and negative percent agreement are, and when each pair applies.

Sensitivity and specificity look like they can be computed from any 2×2 table comparing a new test to an existing one. They cannot, at least not correctly. Both are defined against truth: the true disease status of each subject. If the thing you compared your new test against is not a true reference standard but only another imperfect test, you have not measured sensitivity at all. You have measured agreement. The correct terms are positive and negative percent agreement.

The problem: no gold standard

Sensitivity is the proportion of truly diseased subjects the test flags. To compute it you need to know, independently, who is truly diseased. That means a reference standard: biopsy, culture, long-term follow-up, an established definitive method. Often no such standard is available. The reference method may be too invasive to use on every subject, or it may be imperfect itself. So a new assay gets compared against the best existing test instead.

The moment your comparator is another fallible test rather than the truth, the denominators change meaning. Evaluating a qualitative assay end to end covers where this sits in the wider study. “Of everyone the comparator called positive, how many did the new test also call positive?” is not sensitivity. The comparator’s positives are not the truly diseased. Calling it sensitivity credits the new test with detecting disease when all you have shown is that it agrees with another test that is itself sometimes wrong.

Fourteen subjects shown as three rows of circles. The top row is true disease status, with six diseased. The middle row is the comparator, which misses one diseased subject and wrongly flags one healthy subject. The bottom row is the new test, which matches the comparator exactly, including on both of the subjects the comparator got wrong, which are ringed in the diagram.
A new test can agree perfectly with the comparator and still be wrong about the same subjects. Sensitivity is measured against the top row; percent agreement is measured against the middle one. Same arithmetic, different denominator, different claim.

Positive and negative percent agreement

The correct statistics for this situation are positive percent agreement (PPA) and negative percent agreement (NPA).

PPA = (both tests positive) / (all comparator positives)
NPA = (both tests negative) / (all comparator negatives)

Arithmetically they look just like sensitivity and specificity. Same cell counts, same ratios. But the name is a deliberate signal. The denominator is a comparator method, not a reference standard. No claim about detecting true disease is being made.

The distinction is not pedantry. The US FDA’s guidance on reporting diagnostic test results is explicit. When the comparator is not a reference standard, results should be reported as percent agreement rather than sensitivity and specificity. That way readers are not misled into reading an agreement study as a proof of accuracy. Performance studies for an IVD 510(k) sets out the rest of what that guidance recommends.

Why the wording matters

Two tests can agree beautifully and both be wrong in the same way. High PPA and NPA tell you the new test behaves like the comparator, which is useful if the comparator is trusted and you are seeking equivalence. High agreement says nothing about whether either test is actually right. Reporting that as “98% sensitivity” makes a claim about disease detection the study never supported. In a regulatory submission it is the kind of overstatement that gets the submission rejected.

The wording also changes how a reader should act. Sensitivity permits statements about ruling disease out. Percent agreement permits statements about interchangeability with the comparator. Those are different clinical claims, and the words carry the difference.

When to use which

Your comparator is… Report…
A true reference standard (biopsy, culture, definitive method, confirmed status) Sensitivity and specificity
Another imperfect test or predicate method (no independent truth) Positive and negative percent agreement

The test is simple. Ask whether your comparator establishes the truth or only represents current practice. If you could not rely on the comparator being right in every case, you are measuring agreement. PPA/NPA is what you report.

Reporting with confidence intervals

Each agreement measure is a proportion. Each needs a confidence interval. The Wilson score interval is the sound default, especially with the small positive counts common in early evaluations. When you are comparing the agreement of two methods, or a difference in agreement, the Newcombe interval handles the difference between two proportions properly. As always, a bare percentage without its interval hides how much the sample could be misleading you.

Downloads

Download the CLSI EP12-A2 method comparison example workbook (.xlsx) — two qualitative methods for the same analyte compared, with positive and negative percent agreement and Wilson confidence intervals, ready to open in the Analyse-it trial.

Common mistakes

Reporting sensitivity when you mean PPA. If the denominator is a comparator method rather than confirmed truth, the number is percent agreement. Naming it sensitivity claims something the study cannot support.

Reading high agreement as proof of accuracy. Two tests can agree and both be wrong. Agreement shows equivalence to the comparator, not correctness.

Resolving discordants selectively. Re-testing only the cases where the two disagree, with a third method, and then including those results, biases the estimate. If you resolve discordants, do it by a pre-specified rule and report the unresolved figures too.

Dropping the confidence interval. Early agreement studies often have few positives. The interval on PPA can be wide. Report it.

Report percent agreement with Analyse-it

Analyse-it evaluates qualitative tests per EP12-A2, and the wording follows the comparator, inside Excel:

  • Sensitivity and specificity against a reference standard, or positive and negative percent agreement against a comparator
  • Wilson confidence intervals on each, and a score test for the difference between two tests
  • The same table either way, so the choice is about what you can claim, not about what the software will compute

Every feature from all five editions for 15 days. Qualitative test evaluation is in the Method Validation and Ultimate editions, from US$ 475 a year. Validated against NIST and CLSI reference datasets. See the sensitivity and specificity guide for the reference-standard case; percent agreement counts raw concordance, while Cohen’s kappa corrects it for chance.