Every feature from all five editions for 15 days.
Method Validation edition from US$ 475 a year · 30-day money-back guarantee.
Excel has no built-in ROC curve. A 2×2 table built with COUNTIFS gives the counts at one threshold and stops. No area under the curve, no confidence interval, no comparison between tests, no view of the trade-off across every threshold. A measurement procedure can be precise, linear and well-characterised analytically, and still be clinically useless if it cannot reliably separate positive from negative cases. Set the threshold too low and clinicians see too many false positives; set it too high and disease is missed.
How well does the test discriminate? Is a new biomarker’s AUC better than the established assay’s, or at least within 0.05 of it? Where should the threshold sit when a false positive and a missed case carry different consequences? Analyse-it covers quantitative tests with EP24-A2 ROC analysis and DeLong AUC comparison, qualitative tests per EP12-A2, and compares up to 10 tests at once. The result is the diagnostic accuracy evidence for product labelling, regulatory submissions or publication.
The ROC curve shows how well a quantitative test separates positive from negative cases at every threshold, not just one. Empirical ROC curves for 1 test, up to 10 paired tests or up to 10 independent tests, each plotted against the no-discrimination line. AUC with DeLong confidence intervals and a Z test that the area is better than chance. Predict sensitivity at a fixed false positive fraction, or the false positive fraction at a fixed sensitivity.
A new biomarker is judged against the established assay, not against chance. The DeLong test compares the AUC of paired or independent tests, with equality, equivalence and non-inferiority options. Test, for example, whether a new biomarker’s AUC is within 0.05 of the established assay. Compare up to 10 tests at once, all plotted on the same ROC chart.
The decision plot shows sensitivity, specificity, likelihood ratios, predictive values or cost across every possible threshold. The optimal threshold is chosen by Youden index, by the point closest to (0,1) on the ROC curve or by minimum cost when false positives and false negatives carry different clinical consequences. The bi-histogram of positive and negative outcomes shows the overlap between the two populations.
Sensitivity and specificity at the chosen threshold are the figures a label, a submission or a paper quotes. Sensitivity, specificity, positive and negative predictive values, positive and negative likelihood ratios, diagnostic odds ratio and Youden index, each with confidence intervals. The TP, TN, FP and FN counts behind them are reported too. The dot plot of positive and negative cases shows every result in each group.
A qualitative test gives a positive or negative result, so its performance is a proportion, and the confidence interval on that proportion is the evidence. Sensitivity and specificity with Clopper-Pearson exact or Wilson score confidence intervals, for 1 test, 2 paired tests or 2 independent tests; a mosaic plot shows the outcomes. Between two tests, the difference in sensitivity and specificity with Newcombe confidence intervals and the score Z test. Agreement between two methods with no known true state — positive and negative agreement, kappa and weighted kappa — is part of method comparison.
See diagnostic performance results in detail — ROC curves, AUC comparison, threshold determination and qualitative test evaluation — using CLSI example datasets you can download and follow along with.
2 pages
EP24-A2 — Appendix DDiagnostic performance is one part of the Method Validation edition, alongside measurement system analysis, method comparison and reference intervals.
Related guides in the Learn section: ROC curves and AUC, sensitivity and specificity, PPA and NPA, comparing two tests with DeLong and choosing a cut-off. Further guides cover Cohen’s kappa and weighted kappa, evaluating a qualitative test and performance studies for an IVD 510(k).
ROC curves and diagnostic accuracy are also in the Medical edition, with Bland-Altman agreement, reference intervals and survival analysis. For clinical and biomedical research it is the better fit, from US$ 340 a year. The Method Validation edition adds precision, linearity, detection limits and regression-based method comparison for validating a method.
Try it on your own data first. The 15-day trial is every feature from all five editions — install it and start straight away.
Method Validation edition: US$ 475 per year or US$ 1155 for a perpetual licence. Every purchase carries a 30-day money-back guarantee. Need a quote for purchasing? Add the licence to the cart and save it as a PDF quote.