Diagnostic performance and ROC curve software for method validation ROC curve analysis per EP24-A2 and qualitative test evaluation per EP12-A2 — AUC comparison, optimal threshold determination and diagnostic accuracy metrics.

Every feature from all five editions for 15 days.
Method Validation edition from US$ 475 a year · 30-day money-back guarantee.

Microsoft Excel with the Analyse-it tab selected, showing a diagnostic performance report for OxLDL and LDL: both ROC curves on one chart against the no-discrimination line, the AUC table with DeLong 95% confidence intervals and Z tests, and the Diagnostic Performance task pane open on ROC Curve with the AUC estimator list dropped open. Handwritten notes: Runs inside Excel: every analysis is on the Analyse-it tab; ROC, AUC, decision thresholds, likelihood ratios, predictive values: one Diagnostic Accuracy group; Both ROC curves on one chart; AUC for each test, with its DeLong 95% CI; AUC estimator, and its CI method; The report is an ordinary Excel worksheet: share it, archive it, open it on any PC with Excel.

Establish the diagnostic accuracy of your test

Analyse-it has been a tremendous help. I’ve published and presented at national cardiology meetings and couldn’t have accomplished most of my research without it. Using Analyse-it, I even found errors or omissions in the work of our statistician!
Regina S. Druz, MD, FACC, FASNC
Director, Nuclear Cardiology
North Shore University Hospital

Excel has no built-in ROC curve. A 2×2 table built with COUNTIFS gives the counts at one threshold and stops. No area under the curve, no confidence interval, no comparison between tests, no view of the trade-off across every threshold. A measurement procedure can be precise, linear and well-characterised analytically, and still be clinically useless if it cannot reliably separate positive from negative cases. Set the threshold too low and clinicians see too many false positives; set it too high and disease is missed.

How well does the test discriminate? Is a new biomarker’s AUC better than the established assay’s, or at least within 0.05 of it? Where should the threshold sit when a false positive and a missed case carry different consequences? Analyse-it covers quantitative tests with EP24-A2 ROC analysis and DeLong AUC comparison, qualitative tests per EP12-A2, and compares up to 10 tests at once. The result is the diagnostic accuracy evidence for product labelling, regulatory submissions or publication.

Empirical ROC curves, AUC with DeLong confidence intervals

The ROC curve shows how well a quantitative test separates positive from negative cases at every threshold, not just one. Empirical ROC curves for 1 test, up to 10 paired tests or up to 10 independent tests, each plotted against the no-discrimination line. AUC with DeLong confidence intervals and a Z test that the area is better than chance. Predict sensitivity at a fixed false positive fraction, or the false positive fraction at a fixed sensitivity.

  • 1 test, up to 10 paired tests or up to 10 independent tests/groups
  • Empirical (non-parametric) ROC curves
  • ROC curve with no-discrimination line
  • AUC with DeLong CIs
  • Z test that the AUC is better than chance
  • Predict FPF at fixed sensitivity, sensitivity at fixed FPF or both at fixed threshold
Microsoft Excel showing the ROC Curve section of a diagnostic performance report: ROC curves for OxLDL and LDL, and the AUC table with DeLong 95% confidence intervals, SE, Z statistic and p-value for each test, with the task pane open on ROC Curve. Handwritten notes: ROC curve for each test; AUC with DeLong 95% CI, Z test and p-value; AUC estimator, CI and method.
ROC curves for OxLDL and LDL from EP24-A2 Appendix D, and the AUC of each test with its DeLong 95% confidence interval and Z test against chance.

DeLong comparison of AUC: equality, equivalence, non-inferiority

A new biomarker is judged against the established assay, not against chance. The DeLong test compares the AUC of paired or independent tests, with equality, equivalence and non-inferiority options. Test, for example, whether a new biomarker’s AUC is within 0.05 of the established assay. Compare up to 10 tests at once, all plotted on the same ROC chart.

  • DeLong difference in AUC: equality, equivalence, non-inferiority
  • Comparisons for all pairs of tests or each against a control
  • Overlaid ROC curves for test comparison
Microsoft Excel showing a diagnostic performance report: the AUC table for OxLDL and LDL and beneath it the Comparisons table with the difference in AUC, its DeLong 95% confidence interval, Z statistic and p-value, with the task pane open on Comparisons. Handwritten notes: AUC of each test; Difference in AUC with its 95% CI, Z test and p-value; Equality, equivalence or non-inferiority hypotheses.
The AUC of OxLDL and LDL and the DeLong comparison of the two areas: the difference in AUC, its 95% confidence interval and the Z test for equality.

Decision plot, bi-histogram and optimal threshold by Youden index, closest-to-(0,1) or cost

The decision plot shows sensitivity, specificity, likelihood ratios, predictive values or cost across every possible threshold. The optimal threshold is chosen by Youden index, by the point closest to (0,1) on the ROC curve or by minimum cost when false positives and false negatives carry different clinical consequences. The bi-histogram of positive and negative outcomes shows the overlap between the two populations.

  • Sensitivity vs specificity, likelihood ratios, predictive values or cost
  • Optimal threshold by Youden index, closest-to-(0,1) or cost of diagnosis/misdiagnosis
  • Bi-histogram of positive/negative outcomes
Microsoft Excel showing the Decision Threshold section of a diagnostic performance report for OxLDL: sensitivity and specificity plotted against every threshold, with the task pane open on the Decision section and its plot type, Youden index and optimal threshold options. Handwritten notes: Sensitivity and specificity at every possible threshold; Plot type, and the optimal threshold by Youden or by cost.
Decision plot for OxLDL from EP24-A2 Appendix D: sensitivity and specificity at every threshold.

Sensitivity, specificity, likelihood ratios and predictive values with CIs

Sensitivity and specificity at the chosen threshold are the figures a label, a submission or a paper quotes. Sensitivity, specificity, positive and negative predictive values, positive and negative likelihood ratios, diagnostic odds ratio and Youden index, each with confidence intervals. The TP, TN, FP and FN counts behind them are reported too. The dot plot of positive and negative cases shows every result in each group.

  • Sensitivity / specificity
  • Likelihood ratios
  • Predictive values
  • Odds ratio
  • Youden index
  • TP, TN, FP, FN counts
  • Dot-plot of positive/negative outcomes
Microsoft Excel showing the Sensitivity / Specificity section of a diagnostic performance report for two H. pylori tests: sensitivity, specificity, false positive and false negative proportions with Wilson 95% confidence intervals, the predictive values, and the Comparisons table with Newcombe confidence intervals and a score Z test, with the task pane open on the accuracy options. Handwritten notes: Sensitivity and specificity with Wilson 95% CIs; Predictive values at a prior probability; Comparison of the two tests: difference, CI and Z test; Sensitivity, specificity, LRs, predictive values, odds ratio: tick them.
Sensitivity, specificity, false positive and false negative proportions of the new and old H. pylori tests from EP12-A2, with Wilson 95% confidence intervals. The predictive values and the comparison of the two tests follow.

EP12-A2 qualitative tests: sensitivity, specificity and comparison of two tests

A qualitative test gives a positive or negative result, so its performance is a proportion, and the confidence interval on that proportion is the evidence. Sensitivity and specificity with Clopper-Pearson exact or Wilson score confidence intervals, for 1 test, 2 paired tests or 2 independent tests; a mosaic plot shows the outcomes. Between two tests, the difference in sensitivity and specificity with Newcombe confidence intervals and the score Z test. Agreement between two methods with no known true state — positive and negative agreement, kappa and weighted kappa — is part of method comparison.

  • 1 test, 2 paired tests or 2 independent tests/groups
  • Sensitivity/specificity (Clopper-Pearson exact, Wilson score CIs)
  • Likelihood ratios (Miettinen-Nurminen score CIs)
  • Predictive values (Mercaldo-Wald logit CIs)
  • Predictive values at multiple prior probabilities (prevalences) new in v5.51
  • Difference in sensitivity/specificity (Newcombe score CIs, plus Tango for paired tests and Miettinen-Nurminen for independent tests)
  • Equivalence and non-inferiority tests for sensitivity/specificity new in v5.65
  • McNemar-Mosteller exact, Fisher exact, score Z test
  • Mosaic plot (qualitative)
Microsoft Excel showing a diagnostic performance report for two qualitative H. pylori tests: mosaic plots of each test against the true state, and the 2 x 2 contingency tables with true and false positives and negatives, with the task pane open on Frequencies. Handwritten notes: Mosaic plot of each test against the true state; 2 x 2 table for each test; Contingency table and mosaic plot, ticked here.
The new and old H. pylori qualitative tests from EP12-A2 Example 10.3.1: mosaic plots of each against the true state, and the 2 × 2 tables beneath.

Example analyses

See diagnostic performance results in detail — ROC curves, AUC comparison, threshold determination and qualitative test evaluation — using CLSI example datasets you can download and follow along with.

EP24 A2 Example 1 2 pages EP24-A2 — Appendix D
OxLDL and LDL diagnostic accuracy.
50 subjects, 28 of them with the condition. ROC curves for both markers with AUC, CIs and a test against 0.5 — OxLDL 0.80, LDL 0.56 — and a DeLong comparison of the two curves. A second analysis adds the bi-histogram and decision threshold plot for OxLDL alone.
EP12 A2 Example 1 2 pages EP12-A2 — Example 10.3.1
H. pylori, two qualitative tests against a known state.
102 subjects. Mosaic plots, sensitivity and specificity with Wilson 95% CIs and predictive values at the observed prior. The difference between the two tests with Newcombe CIs and a score Z test.
EP12 A2 Example 2 1 page EP12-A2 — Example 10.3.2
H. pylori, agreement without a reference method.
536 samples compared between two qualitative methods with no true state available. Positive and negative agreement with Wilson CIs, overall agreement and Cohen’s kappa with a Wald 95% CI.

Part of the Method Validation edition

Diagnostic performance is one part of the Method Validation edition, alongside measurement system analysis, method comparison and reference intervals.

Related guides in the Learn section: ROC curves and AUC, sensitivity and specificity, PPA and NPA, comparing two tests with DeLong and choosing a cut-off. Further guides cover Cohen’s kappa and weighted kappa, evaluating a qualitative test and performance studies for an IVD 510(k).

ROC curves and diagnostic accuracy are also in the Medical edition, with Bland-Altman agreement, reference intervals and survival analysis. For clinical and biomedical research it is the better fit, from US$ 340 a year. The Method Validation edition adds precision, linearity, detection limits and regression-based method comparison for validating a method.

Software you can trust

Validated calculations you can defend at inspection Every calculation is performed by Analyse-it — no Excel formulas, no third-party functions. Results are validated against CLSI reference datasets, published datasets and thousands of internal test cases before every release. How Analyse-it is developed and validated →
Data stays in your facility Analyse-it runs entirely within Microsoft Excel on your PC. No cloud processing, no data transmission. Pre-submission data, data derived from patients and in-process results stay within your facility under your own data governance controls.
Standard Excel workbooks anyone can open Every analysis is an ordinary .xlsx workbook. Share with colleagues, submit to regulatory affairs, archive for audit, open on any PC with Excel. No proprietary format, no licence required to view results. Colleagues and auditors see exactly what you see.
Results that cannot be accidentally broken Analysis output contains computed values, not formulas. Nothing to accidentally overwrite, no cell references to break, no formula errors to introduce. The results you reported are exactly what you will find when you reopen the workbook months or years later for an audit.

Free trial and pricing

Try it on your own data first. The 15-day trial is every feature from all five editions — install it and start straight away.

Method Validation edition: US$ 475 per year or US$ 1155 for a perpetual licence. Every purchase carries a 30-day money-back guarantee. Need a quote for purchasing? Add the licence to the cart and save it as a PDF quote.