Evaluating a qualitative test A test that returns positive or negative still has precision, a measuring range, and accuracy — they just take a different form. How to evaluate a qualitative assay end to end under CLSI EP12-A2.

A qualitative test reports a category (positive or negative, reactive or non-reactive) rather than a number. That does not make it simpler to evaluate. The evaluation simply takes an unfamiliar shape.

Accuracy still has to be measured, and so does precision. A binary result has no standard deviation, so the usual precision statistics do not apply. CLSI EP12 sets out how to do it. The whole approach depends on one region: the concentrations near the cut-off, where the test switches from calling samples negative to calling them positive.

A note on editions before the detail. The current edition is EP12-Ed3, Evaluation of Qualitative, Binary Output Examination Performance, which replaced EP12-A2, User Protocol for Evaluation of Qualitative Test Performance. The change of title reflects a change of scope: the third edition covers evaluation by the manufacturer as well as by the user, and CLSI publishes a separate implementation guide, EP12IG, for the user verification case. The worked examples on this page are numbered from EP12-A2, but the statistics below are common to both editions.

Design the study around the cut-off

Far above the cut-off, a decent qualitative test calls everything positive. Far below it, everything negative. Neither region tells you much. All the interesting behaviour (the disagreements, the run-to-run inconsistency) lives in the transition zone straddling the cut-off. An EP12 study concentrates its samples there, with enough spread on either side to map where the test flips. Loading up on obvious positives and obvious negatives inflates your agreement figures while telling you nothing about the assay’s real weak point.

Accuracy: against a reference or a comparator

How you express accuracy depends on what you compared against. The distinction matters enough to have its own guide. If you have a true reference standard (confirmed disease status, a definitive method), you report sensitivity and specificity. If your comparator is only another imperfect test, you cannot claim sensitivity. You report positive and negative percent agreement instead. Choosing the wrong pair of terms overstates what the study showed. Either way, report each figure with a Wilson confidence interval, because the counts near the cut-off are usually small.

Precision: the C5–C95 interval

A qualitative test’s precision is the width of its transition zone. Define the C5 concentration as the one at which the test calls samples positive 5% of the time, and C95 as the one where it does so 95% of the time.

The interval between them, the C5–C95 interval, is the concentration range over which the result is genuinely uncertain. That interval is the qualitative analogue of imprecision. A narrow interval means the test flips crisply and reproducibly at its cut-off. A wide one means results in that band are close to random.

You estimate the interval by testing samples at several concentrations across the transition, many replicates each, and modelling the proportion positive against concentration. A narrow C5–C95 interval is what lets you trust a near-cut-off result. A wide one is a warning to interpret borderline results cautiously. The width is worth comparing across reagent lots, instruments, and operators.

An S-shaped curve of the probability of a positive call against concentration, rising from near zero to near one, with dashed lines at 5% and 95% marking C5 and C95 and the interval between them shaded as the grey zone.
A qualitative test’s precision is the width of its grey zone. As concentration rises, the probability of a positive call climbs from near zero to near one. C5 and C95 mark where it crosses 5% and 95%. The interval between them is the range over which the result is unreliable.

Reproducibility across conditions

The transition zone is where a qualitative test is fragile. That is where reproducibility should be checked: does the C5–C95 interval hold across operators, instruments, reagent lots, and sites? A test that flips crisply for one operator and sloppily for another has a reproducibility problem that a single-operator study would never reveal. For an assay heading into multi-site use, agreement of the transition zone across sites is as important as the reported sensitivity.

Downloads

Download the CLSI EP12-A2 qualitative evaluation example workbook (.xlsx) — a new test evaluated against a comparator, with sensitivity, specificity, and agreement reported with Wilson confidence intervals, ready to open in the Analyse-it trial.

Mosaic plots and frequency tables for two qualitative tests each evaluated against a known true state, followed by a comparison table giving the difference in sensitivity and specificity with Newcombe confidence intervals, and a score Z test for each difference.
Two qualitative methods for H. pylori, each evaluated against a known true state (CLSI EP12-A2, Example 10.3.1). The comparison below gives the difference between them: sensitivity differs by 0.049 with a Newcombe interval spanning zero, so that difference is not established; specificity differs by 0.122 with an interval clear of zero, and the score test agrees (p = 0.0253).

Common mistakes

Testing away from the cut-off. Obvious positives and negatives pad the agreement statistics and hide the transition zone where the test actually struggles. Concentrate samples near the cut-off.

Reporting sensitivity against a non-reference comparator. If the comparator is not a true standard, the figure is percent agreement, not sensitivity. Name it correctly.

Ignoring qualitative precision entirely. “It’s just positive or negative” is not a reason to skip precision. The C5–C95 interval is the precision, and near-cut-off results depend on it.

Dropping confidence intervals. Near-cut-off counts are small and the intervals are wide. A bare percentage overstates how well the test is pinned down.

Evaluate a qualitative test with Analyse-it

Analyse-it runs the EP12-A2 evaluation on your own results, inside Excel:

  • Sensitivity and specificity against a reference standard, or positive and negative percent agreement against a comparator
  • Wilson confidence intervals on every proportion, including the ones near the boundary
  • The C5–C95 interval, which is what qualitative precision means

Every feature from all five editions for 15 days. Qualitative test evaluation is in the Method Validation and Ultimate editions, from US$ 475 a year. Validated against NIST and CLSI reference datasets. See also sensitivity and specificity and Cohen’s kappa for agreement corrected for chance.