DeLong's paired test compares two ROC curves measured on the same patients — the standard way to decide whether one diagnostic test genuinely outperforms another.
DeLong's method derives the variance and covariance of two paired AUCs from their Mann-Whitney U representation, giving a z-test for the AUC difference. Because the test is paired (both predictors measured on the same patients), it has much more power than comparing two independent AUC estimates.
Optional partial-AUC support restricts the comparison to a clinically relevant region of the ROC (e.g. only high-specificity territory, FPR < 0.1) and reports the McClish-corrected partial AUC for both predictors.
Binary label + two numeric predictors paired by row.
Both AUCs, AUC difference + CI, z, p (one- or two-sided). With partial-AUC mode: McClish-corrected partial AUCs over a specificity range.
A significant test with a small ΔAUC (e.g. 0.01) is rarely clinically meaningful — magnitude over significance.
Sample size matters: with n > 1000 per arm, tiny ΔAUCs become 'significant'; report the CI.