Guides by field

Data scientists & ML practitioners

Predict, cluster, reduce dimensions, and validate — the machine-learning toolkit.

Machine learning is prediction under a different name, and the same discipline applies: match the model to the target, hold out data to estimate honest performance, and prefer the simplest model that does the job. This guide maps the supervised, unsupervised, and evaluation tools onto the questions you actually ask.

A running theme — interpretable models (logistic regression, a single tree) buy you explanation; ensembles and kernels (random forest, boosting, SVM) usually buy you accuracy. Start simple, then add complexity only if it earns its keep on held-out data.

The non-negotiable discipline is honest evaluation: estimate performance on data the model never saw — a held-out test set or k-fold cross-validation — because training accuracy always flatters. Standardize features before any distance- or gradient-based method (kNN, k-means, SVM, PCA, neural nets), and when classes are imbalanced prefer precision-recall / F1 over plain accuracy and ROC-AUC, which can look strong while missing the rare positives.

Predicting a category (classification) (ISL Ch. 4, 8–9)

When the target is a label — churn / no-churn, disease class, spam — you have plenty of choices. Logistic regression and a single decision tree are transparent and a strong baseline; random forest and gradient boosting usually win on raw accuracy; SVM shines with many features; kNN and naive Bayes are fast, simple yardsticks.

  • — Interpretable baseline for a binary target (coefficients = log-odds).
  • — Transparent if-then rules; the building block of ensembles.
  • — Accurate, robust ensemble; gives feature importance.
  • — Gradient-boosted trees — often the top accuracy on tabular data.
  • — Margin-based classifier, strong in high dimensions.
  • — Simple instance-based baseline.
  • — Three or more unordered classes, with interpretable coefficients.
  • — Fast probabilistic baseline (text-friendly).
  • — Linear classifier that doubles as dimension reduction.

Predicting a number & taming many predictors (ISL Ch. 3, 6)

For a numeric target, start with multiple regression; when predictors are many or collinear, regularization (lasso / ridge) prevents overfitting — and lasso also selects features. Quantile regression models the tails, not just the mean; a neural net is the flexible non-linear fallback.

  • — The linear baseline for a numeric target.
  • — Lasso / ridge — shrink coefficients and select features.
  • — Many correlated predictors reduced to latent components that predict Y.
  • — Custom family + link when the outcome is neither Gaussian nor binary.
  • — Model the 10th/90th percentile, not just the average.
  • — Flexible non-linear function approximator.

Reducing dimensions & visualizing (ISL Ch. 12)

High-dimensional tables are hard to see and to model. PCA gives interpretable linear components; t-SNE and UMAP produce 2D embeddings that reveal clusters (UMAP keeps more global structure and runs faster); MDS lays points out to match a distance matrix.

  • — Interpretable linear components; a modeling and viz workhorse.
  • — Fast non-linear 2D embedding, preserves global structure.
  • — Non-linear embedding tuned for local cluster separation.
  • — Non-metric MDS when only the rank order of distances is meaningful.
  • — Shared structure between two blocks of variables.
  • — Lay points out to reproduce a distance/dissimilarity matrix.
  • — The categorical analogue — biplot of a contingency table.

Finding groups (clustering) (ISL Ch. 12)

Unsupervised structure discovery. k-means is the fast default (use the elbow to pick k); DBSCAN finds arbitrary shapes and flags noise; hierarchical clustering gives a dendrogram so you choose the cut after seeing the structure.

  • — Fast partitioning into k spherical clusters.
  • — Pick k for k-means from the within-cluster scree.
  • — Density-based — arbitrary shapes + noise points.
  • — Dendrogram; choose the number of clusters visually.
  • — Fuzzy c-means — soft (probabilistic) membership.

Evaluating a model honestly (ISL Ch. 5)

Training accuracy is a mirage. Judge classifiers on held-out data with ROC/AUC and precision-recall (better when positives are rare), compare two models’ AUCs directly, and use resampling to put confidence intervals on any statistic.

  • — AUC + threshold selection for a scoring model.
  • — Is model A’s AUC genuinely higher than B’s? (DeLong).
  • — Precision-recall — the honest view under class imbalance.
  • — Confidence intervals for any metric via resampling.
  • — Screen many features at once, with a volcano plot and multiplicity control.
  • — Distribution-free significance test.

References & further reading

Other fields: · · · · · · · ·