Relate › Correlation

Pearson's r

Pearson's product-moment correlation measures the strength of linear association between two continuous variables, running from −1 (perfect inverse) through 0 (none) to +1 (perfect positive).

What is Pearson's r?

Pearson's r is the standardised covariance — the covariance divided by the product of the SDs of the two variables. By construction it captures only the *linear* component of the relationship; a perfect parabolic relationship (y = x²) can have r = 0 even though x and y are perfectly determined.

Inference (the p-value and CI) is calibrated under the assumption of bivariate normality. With non-normal data the point estimate of r is still meaningful, but the p-value can be wrong. For ordinal data, monotonic-non-linear relationships, or non-normal continuous data, switch to Spearman's ρ or Kendall's τ — both are rank-based and make no distributional assumptions for inference.

Pearson is highly sensitive to outliers — a single extreme point can pull r from 0 to 0.9 or vice versa. Always plot before computing; if outliers are present and the rest of the data is clearly correlated, rank-based correlation (Spearman / Kendall) gives the more honest answer.

When should I use Pearson's r?

  • Both variables continuous, approximately normal, and the relationship is linear.
  • Sanity-checking covariance / SDs from a clean dataset before linear regression.
  • Switch to Spearman when normality is questionable or the relationship is monotonic-but-curved.
  • Switch to Kendall when ties dominate or n is small (n < 30).

What data does it need?

Two numeric columns.

What does it report?

r, t statistic, df = n − 2, two-sided p, 95% CI via Fisher's z-transform.

What does it assume?

  • Independent paired observations.
  • Both variables interval/ratio (truly continuous).
  • Approximately bivariate normal.
  • Linear relationship.
  • No extreme outliers.

Formula

r = Σ((xᵢ − x̄)(yᵢ − ȳ)) / √(Σ(xᵢ − x̄)² · Σ(yᵢ − ȳ)²)
Fisher z = ½ · log((1 + r) / (1 − r))

How do I interpret the result?

r² is the proportion of variance in one variable linearly explained by the other; r = 0.7 means 49% shared variance, not 70%.

Cohen anchors: |r| ≈ 0.1 small, 0.3 medium, 0.5 large — calibration, not law.

A statistically significant r at large n can still be tiny; the CI tells you whether the magnitude is meaningful.

See also

References

  • Pearson (1895). Note on regression and inheritance in the case of two parents. Proc. Royal Soc. London 58.