Categorical › Classification (supervised)

k-nearest neighbours (kNN)

k-nearest neighbours classifies each observation by majority vote of its k closest standardized neighbours, validated leave-one-out.

What is k-nearest neighbours (kNN)?

kNN is a non-parametric, instance-based classifier: no model is fit — a point's class is decided by the classes of its k nearest points in predictor space. Small k tracks local structure (low bias, high variance); large k smooths (vice versa).

Because distances drive everything, standardization matters: an unstandardized predictor with a large scale dominates. LOOCV classifies each point with itself held out — the reported accuracy is honest.

When should I use k-nearest neighbours (kNN)?

  • Quick nonlinear classification baseline.
  • Decision boundaries too irregular for linear methods.

What data does it need?

Categorical outcome + numeric predictors + k + standardization toggle.

What does it report?

LOOCV accuracy, confusion matrix, per-class error.

What does it assume?

  • Meaningful distance metric (numeric predictors, sensibly scaled).
  • Enough data density near class boundaries.

How do I interpret the result?

Try a few k values; accuracy that collapses with small k suggests noise, with large k suggests over-smoothing.

See also

References

  • Hastie, Tibshirani & Friedman (2009). Elements of Statistical Learning, ch. 13.