Categorical › Classification (supervised)

Random forest

A random forest classifies or predicts by averaging an ensemble of decorrelated decision trees, reporting out-of-bag error and variable importance.

What is Random forest?

Each tree is grown on a bootstrap sample using a random subset of predictors per split; averaging many decorrelated trees slashes variance. Observations not in a tree's bootstrap sample (out-of-bag) provide honest error estimates without a holdout set.

Outcome type is auto-detected: categorical → classification (OOB accuracy + confusion), numeric → regression (OOB %variance explained + MSE).

Importance: mean decrease in accuracy (permutation-based, more reliable) and mean decrease in Gini / node purity (fit-based, fast).

When should I use Random forest?

  • Strong tabular-data baseline with minimal tuning.
  • Ranking predictor importance.
  • Nonlinear relationships and interactions without specifying them.

What data does it need?

Outcome (categorical or numeric) + numeric predictors + number of trees.

What does it report?

OOB accuracy/confusion (classification) or %Var/MSE (regression) + sorted variable-importance table.

What does it assume?

  • Independent observations.
  • OOB estimates assume the forest has enough trees to stabilize.

How do I interpret the result?

OOB accuracy is comparable to cross-validation. Importance ranks predictors; permutation importance (MDA) is the safer of the two columns.

See also

References

  • Breiman (2001). Random Forests. Machine Learning 45.