Diagnostics › Diagnostic accuracy

Compare two AUCs (DeLong's test)

DeLong's paired test compares two ROC curves measured on the same patients — the standard way to decide whether one diagnostic test genuinely outperforms another.

What is Compare two AUCs (DeLong's test)?

DeLong's method derives the variance and covariance of two paired AUCs from their Mann-Whitney U representation, giving a z-test for the AUC difference. Because the test is paired (both predictors measured on the same patients), it has much more power than comparing two independent AUC estimates.

Optional partial-AUC support restricts the comparison to a clinically relevant region of the ROC (e.g. only high-specificity territory, FPR < 0.1) and reports the McClish-corrected partial AUC for both predictors.

When should I use Compare two AUCs (DeLong's test)?

  • Comparing two diagnostic tests measured on the same patients.
  • Head-to-head comparison of two model scores on the same validation set.
  • Showing a new biomarker beats the standard one.

What data does it need?

Binary label + two numeric predictors paired by row.

What does it report?

Both AUCs, AUC difference + CI, z, p (one- or two-sided). With partial-AUC mode: McClish-corrected partial AUCs over a specificity range.

What does it assume?

  • Independent patients (each row is one subject).
  • Both scores measured on every subject (no missing pairs).
  • Correct binary gold standard.

How do I interpret the result?

A significant test with a small ΔAUC (e.g. 0.01) is rarely clinically meaningful — magnitude over significance.

Sample size matters: with n > 1000 per arm, tiny ΔAUCs become 'significant'; report the CI.

See also

References

  • DeLong, DeLong & Clarke-Pearson (1988). Comparing the areas under two or more correlated ROC curves: a nonparametric approach. Biometrics 44(3).