An ROC curve shows how well a continuous biomarker separates a binary outcome across every possible cutoff.
Each point on the ROC curve is a (1 − specificity, sensitivity) pair for one decision threshold of a continuous test. As the threshold moves from very low (everyone positive) to very high (everyone negative), the point sweeps from (1, 1) down to (0, 0). AUC = area under that curve = the probability that the test ranks a random positive subject higher than a random negative.
AUC is threshold-free, so it's the standard summary of a test's discrimination across all possible cutoffs. AUC of 0.5 ⇒ no better than chance; 0.7 = fair; 0.8 = good; 0.9 = excellent (clinical convention).
ROC is for *discrimination* — how well the test ranks positives above negatives. Calibration (do the predicted probabilities match observed event rates) is a separate question, addressed by reliability plots or the Hosmer-Lemeshow test.
Numeric predictor + binary label + optional grouping factor.
ROC curve, AUC + CI, Youden-optimal threshold. Interactive panel computes sensitivity / specificity / PPV / NPV / LR+ / LR− / accuracy / F1 / MCC / κ at any user-selected threshold; PPV / NPV recompute under user-supplied population prevalence.
Youden J = sensitivity + specificity − 1 maximises the (1, 1) − (FPR, TPR) Pythagoras distance. It picks the 'fairest' cutoff but ignores asymmetric misclassification costs.
PPV and NPV are prevalence-dependent — the same test has very different predictive values in screening (low prevalence) vs diagnostic (high prevalence) settings. Always report the prevalence used.
On highly imbalanced data, AUC can be misleading; consider AUPRC instead (precision-recall curve area).