Paired comparison of two precision-recall curves measured on the same patients tests the difference with a percentile bootstrap.
Precision-recall curves are more informative than ROC curves for highly imbalanced binary outcomes (e.g. fraud detection, rare disease screening) because the FPR axis in ROC compresses the meaningful action region into a tiny range. AUPRC captures the trade-off you actually care about: precision (PPV) at varying recall (sensitivity).
Comparing two AUPRCs on the same patients (paired data) preserves the within-subject correlation in the bootstrap. The percentile CI on Δ AUPRC has no closed-form analogue to the DeLong test for ROC — bootstrap is the standard.
Outcome (binary) + predictor A + predictor B (paired) + # bootstrap reps.
AUPRC per predictor + Δ AUPRC + percentile bootstrap CI; PR curves on one plot with random-classifier baseline.
The baseline AUPRC for a random classifier equals the prevalence of positives — judge improvements over this baseline, not over 0.5.