Categorical › Classification (supervised)

Boosting (gradient boosted trees)

Gradient boosting fits an ensemble of shallow trees sequentially to the residuals, choosing the stopping iteration by 5-fold cross-validation, for classification or regression.

What is Boosting (gradient boosted trees)?

Boosting builds trees one at a time, each fitting what the previous ensemble got wrong, shrunk by the learning rate. Small learning rates + more trees generalize better; interaction depth bounds how many variables can interact within one tree.

The CV-selected best iteration guards against overfitting — performance is reported from cross-validated predictions at that iteration, not the training fit. Relative influence measures each predictor's total contribution to loss reduction.

When should I use Boosting (gradient boosted trees)?

  • Strong tabular baseline when random forest underfits smooth signals.
  • When you want variable importance from an additive-tree model.

What data does it need?

Outcome (categorical or numeric) + numeric predictors + max trees, interaction depth, shrinkage.

What does it report?

CV accuracy + CV confusion (classification) or CV RMSE + R² (regression), best iteration, relative-influence table.

What does it assume?

  • Independent observations.
  • Needs ≥ ~40 rows for 5-fold CV with tree minimums.

How do I interpret the result?

If the best iteration equals the max trees, raise the cap — the model wanted to keep learning. Compare CV metrics to random forest; boosting usually wins on smooth effects, forests on messy interactions.

See also

References

  • Friedman (2001). Greedy function approximation: a gradient boosting machine. Ann Statist 29.