Diagnostics › Diagnostic accuracy

Compare two precision-recall curves (bootstrap)

Paired comparison of two precision-recall curves measured on the same patients tests the difference with a percentile bootstrap.

What is Compare two precision-recall curves (bootstrap)?

Precision-recall curves are more informative than ROC curves for highly imbalanced binary outcomes (e.g. fraud detection, rare disease screening) because the FPR axis in ROC compresses the meaningful action region into a tiny range. AUPRC captures the trade-off you actually care about: precision (PPV) at varying recall (sensitivity).

Comparing two AUPRCs on the same patients (paired data) preserves the within-subject correlation in the bootstrap. The percentile CI on Δ AUPRC has no closed-form analogue to the DeLong test for ROC — bootstrap is the standard.

When should I use Compare two precision-recall curves (bootstrap)?

  • Comparing classifiers in heavily imbalanced settings where AUPRC is more meaningful than AUROC.
  • Head-to-head model comparison on the same evaluation cohort.

What data does it need?

Outcome (binary) + predictor A + predictor B (paired) + # bootstrap reps.

What does it report?

AUPRC per predictor + Δ AUPRC + percentile bootstrap CI; PR curves on one plot with random-classifier baseline.

How do I interpret the result?

The baseline AUPRC for a random classifier equals the prevalence of positives — judge improvements over this baseline, not over 0.5.

See also

References

  • Davis & Goadrich (2006). The relationship between Precision-Recall and ROC curves. ICML.