Pearson's product-moment correlation measures the strength of linear association between two continuous variables, running from −1 (perfect inverse) through 0 (none) to +1 (perfect positive).
Pearson's r is the standardised covariance — the covariance divided by the product of the SDs of the two variables. By construction it captures only the *linear* component of the relationship; a perfect parabolic relationship (y = x²) can have r = 0 even though x and y are perfectly determined.
Inference (the p-value and CI) is calibrated under the assumption of bivariate normality. With non-normal data the point estimate of r is still meaningful, but the p-value can be wrong. For ordinal data, monotonic-non-linear relationships, or non-normal continuous data, switch to Spearman's ρ or Kendall's τ — both are rank-based and make no distributional assumptions for inference.
Pearson is highly sensitive to outliers — a single extreme point can pull r from 0 to 0.9 or vice versa. Always plot before computing; if outliers are present and the rest of the data is clearly correlated, rank-based correlation (Spearman / Kendall) gives the more honest answer.
Two numeric columns.
r, t statistic, df = n − 2, two-sided p, 95% CI via Fisher's z-transform.
r² is the proportion of variance in one variable linearly explained by the other; r = 0.7 means 49% shared variance, not 70%.
Cohen anchors: |r| ≈ 0.1 small, 0.3 medium, 0.5 large — calibration, not law.
A statistically significant r at large n can still be tiny; the CI tells you whether the magnitude is meaningful.