Every feature from all five editions for 15 days, no sign-up.
Medical edition from US$ 340 a year · 30-day money-back guarantee.
Excel has no ROC curve. A 2×2 table built with COUNTIFS gives the counts at one threshold and stops. No area under the curve, no confidence interval, no comparison between tests, no view of how sensitivity and specificity trade off as the threshold moves. A diagnostic test is only useful if it discriminates reliably between patients with and without the condition. A test that looks adequate on AUC alone may fail at the threshold that matters clinically.
How well does the test discriminate? Does a new biomarker outperform the test clinicians already use, or match it closely enough to replace it? Where should the threshold sit when a missed diagnosis costs more than a false positive? Analyse-it evaluates quantitative tests with ROC curves and qualitative tests against the true state, and compares up to 10 tests at once. The result is the sensitivity, specificity and predictive values a screening, staging or referral decision needs.
The ROC curve shows how well a quantitative test separates patients with the condition from those without at every threshold, not just one. Empirical ROC curves for a single test, up to 10 paired tests or up to 10 independent tests or groups, each plotted against the no-discrimination line. Wilcoxon-Mann-Whitney AUC with DeLong-DeLong-Clarke-Pearson confidence intervals and a Z test that the area is better than chance. Predict the false positive fraction at a fixed sensitivity, sensitivity at a fixed false positive fraction or both at a fixed threshold.
A new test is judged against the one clinicians already use, not against chance. Overlay the ROC curves of up to 10 tests on one plot and compare their AUCs with the DeLong test. Test for equality, for equivalence or for non-inferiority when the new test only needs to match the established one. Compare every pair of tests, or each new test against the established one, with the difference in AUC and its confidence interval alongside the decision.
Set the threshold too low and the clinic sees too many false positives; set it too high and disease is missed. The decision plot shows sensitivity and specificity, likelihood ratios, predictive values or cost across every possible threshold, so the trade-off is visible before a cut-off is chosen. The optimal threshold is chosen by Youden index, by the point closest to (0,1) on the ROC curve or by cost. Weight a missed cancer diagnosis at 10× the cost of a false positive biopsy referral, for example.
Sensitivity and specificity at the chosen threshold are the figures a paper, a guideline or a screening protocol quotes. Sensitivity and specificity with Clopper-Pearson exact or Wilson score confidence intervals, positive and negative likelihood ratios, diagnostic odds ratio and Youden index. Positive and negative predictive values at any prevalence — 5% for population screening, say, against 40% in a specialist referral clinic. The TP, TN, FP and FN counts behind them are reported, and the bi-histogram and dot plot show the separation between positive and negative cases.
When the result is positive or negative rather than a value, performance is a proportion, and the confidence interval on that proportion is the evidence. Sensitivity and specificity with Clopper-Pearson exact or Wilson score intervals, likelihood ratios, predictive values, diagnostic odds ratio and Youden index. One test, two paired tests or two independent tests or groups. Agreement between two methods with no known true state — PPA and NPA, kappa and weighted kappa — is part of method agreement. Compare sensitivity and specificity between two tests with Newcombe, Tango or Miettinen-Nurminen score intervals and the McNemar-Mosteller exact, Fisher exact or score Z test. A mosaic plot shows the outcomes.
See diagnostic accuracy results in detail — ROC curves, AUC comparison, decision plots and qualitative test evaluation — using example datasets you can download and follow along with.
2 pages
EP24-A2 — Appendix DDiagnostic accuracy evaluation is one part of the Medical edition, alongside Bland-Altman agreement, reference intervals and survival analysis. The edition also includes the full Standard edition for hypothesis testing, regression and descriptive statistics.
Related guides in the Learn section: ROC curves and AUC, sensitivity, specificity and predictive values, likelihood ratios, and PPA and NPA. Further guides cover choosing a cut-off, comparing two tests with DeLong and performance studies for an IVD 510(k).
For the rest of a method validation programme — precision, linearity, detection limits, bias at clinical decision points, regression-based method comparison and reference-interval transference — see the Method Validation edition.
Try it on your own data first. The 15-day trial is every feature from all five editions, with no sign-up and no licence key — install it and start straight away.
Medical edition: US$ 340 per year or US$ 815 for a perpetual licence. Every purchase carries a 30-day money-back guarantee. Need a quote for purchasing? Add the licence to the cart and save it as a PDF quote.