Diagnostics › Diagnostic accuracy

ROC (binary)

An ROC curve shows how well a continuous biomarker separates a binary outcome across every possible cutoff.

What is ROC (binary)?

Each point on the ROC curve is a (1 − specificity, sensitivity) pair for one decision threshold of a continuous test. As the threshold moves from very low (everyone positive) to very high (everyone negative), the point sweeps from (1, 1) down to (0, 0). AUC = area under that curve = the probability that the test ranks a random positive subject higher than a random negative.

AUC is threshold-free, so it's the standard summary of a test's discrimination across all possible cutoffs. AUC of 0.5 ⇒ no better than chance; 0.7 = fair; 0.8 = good; 0.9 = excellent (clinical convention).

ROC is for *discrimination* — how well the test ranks positives above negatives. Calibration (do the predicted probabilities match observed event rates) is a separate question, addressed by reliability plots or the Hosmer-Lemeshow test.

When should I use ROC (binary)?

  • Evaluating a continuous biomarker / score against a binary outcome.
  • Choosing a clinical cutoff via Youden's J or a fixed-specificity / fixed-sensitivity rule.
  • Reporting a model's discrimination — pair with calibration analysis for the full picture.

What data does it need?

Numeric predictor + binary label + optional grouping factor.

What does it report?

ROC curve, AUC + CI, Youden-optimal threshold. Interactive panel computes sensitivity / specificity / PPV / NPV / LR+ / LR− / accuracy / F1 / MCC / κ at any user-selected threshold; PPV / NPV recompute under user-supplied population prevalence.

What does it assume?

  • Independent observations.
  • Binary, correctly classified gold standard.
  • Continuous predictor (or at least an ordinal score).

Formula

AUC = P(score_positive > score_negative) (Mann-Whitney U interpretation)

How do I interpret the result?

Youden J = sensitivity + specificity − 1 maximises the (1, 1) − (FPR, TPR) Pythagoras distance. It picks the 'fairest' cutoff but ignores asymmetric misclassification costs.

PPV and NPV are prevalence-dependent — the same test has very different predictive values in screening (low prevalence) vs diagnostic (high prevalence) settings. Always report the prevalence used.

On highly imbalanced data, AUC can be misleading; consider AUPRC instead (precision-recall curve area).

See also

References

  • Hanley & McNeil (1982). The meaning and use of the area under a receiver operating characteristic (ROC) curve. Radiology 143(1).