Guides by field

Starting out — describe & explore

Look at the data, check its shape, and answer the simplest questions — before any modelling.

Every analysis in every field starts the same way: look at the data. This guide is the one to read first, whatever your field — it covers describing a dataset, plotting it, checking the assumptions the fancier tests rely on, and answering the simplest questions (is this mean different from a target? do these two things move together?).

The order matters. Plot before you test: a histogram or box plot catches the skew, the bimodality, and the stray decimal point that no p-value will tell you about. Then check the assumptions — most classical tests want roughly normal data — and only then reach for the test. If an assumption fails you usually have a rank-based alternative a click away.

One habit worth forming early: decide what you are asking before you compute. A test answers "could chance alone produce a difference this big?", which is not the same as "is this difference big enough to matter". Report an effect size and a confidence interval alongside every p-value, and let the interval — not the asterisk — carry the conclusion.

Describing a dataset

Start with a summary panel across every numeric column: n, missing, mean, SD, median, quartiles, and the shape statistics (skewness and kurtosis). It is the fastest way to spot a column that is mostly empty, a unit mix-up, or an impossible minimum.

  • — n, missing, mean, SD, CI, median, IQR, skewness, kurtosis — all columns at once.
  • — Directions, angles, or times of day, where 359° and 1° are neighbours.

Seeing the distribution

A histogram shows the shape; a box plot compares groups compactly; a violin adds the density the box hides; a raincloud shows shape, summary, and every raw point at once. If you are comparing two groups and want the effect size to be the headline rather than the p-value, the estimation plot puts the mean difference and its confidence interval next to the raw data.

  • — Shape of one column — skew, gaps, multiple peaks.
  • — Median, quartiles and outliers, side by side across groups.
  • — The density a box plot hides (bimodality especially).
  • — Density + box + every raw point — the most honest single view.
  • — Small samples, where every observation deserves to be visible.
  • — Two groups with the mean difference and its CI as the headline.

Is the data normal enough?

Most classical tests assume roughly normal data. Judge that with a Q-Q plot first — it shows you *how* the data departs from normal, which a single p-value cannot. Use a formal test to back up the picture, remembering that with a large n these tests flag trivial departures, and with a small n they miss real ones.

  • — The first thing to look at: points on the line mean normal.
  • — Companion to the Q-Q plot, more sensitive in the middle of the distribution.
  • — The default normality test — good power at small to moderate n.
  • — Omnibus test combining skewness and kurtosis.
  • — Anderson-Darling — more weight on the tails.
  • — Kolmogorov-Smirnov (Lilliefors-corrected).
  • — Lilliefors — the correct null when μ and σ come from the sample.
  • — Shapiro-Francia — a simpler variant that suits larger samples.

Outliers: real signal or data-entry slip?

Investigate an outlier before deleting it — a value can be extreme and completely correct. These rules identify candidates; the decision to exclude is yours, and belongs in your write-up.

  • — Tukey's fence: outside Q1 − 1.5·IQR or Q3 + 1.5·IQR.
  • — Dixon's Q — built for very small samples (n ≤ 30).

Do two variables move together?

Correlation measures association, never causation — and Pearson measures *linear* association specifically, so plot the scatter first. If the relationship is monotone but curved, or the data is ordinal or outlier-prone, the rank-based correlations are the honest choice. Partial correlation asks the same question while holding other variables fixed.

  • — Always look first — correlation cannot see a curve or a cluster.
  • — Pearson's r: linear association between two continuous variables.
  • — Spearman's ρ: monotone association, robust to outliers and skew.
  • — Kendall's τ: concordance-based, good for small n and many ties.
  • — Association between X and Y with control variables held fixed.
  • — Adds a third variable as point size.

Comparing against a known value

The simplest hypothesis of all: does this sample differ from a number you already have — a target, a specification, a published rate? Pick by outcome type: a mean, a median, or a proportion.

  • — Is the mean different from a target value?
  • — The same question for a median, without assuming normality.
  • — Is an observed proportion different from a hypothesised one? (exact).
  • — Do counts across several categories match expected proportions?

Charts for particular shapes of data

Some data has a shape that a scatter or bar chart cannot show honestly: compositions that sum to a whole, angles, running totals, or a response surface over two inputs. These plots exist for those cases.

  • — How a total is built up (or eaten away) step by step.
  • — Several measures per item, compared across items.
  • — Three-part compositions that sum to 100%.
  • — Angle-and-radius data (wind, time of day, orientation).
  • — A Z surface over an X-Y plane, drawn as iso-value contours.
  • — Plot a typed formula y = f(x) — handy for overlaying theory.
  • — Paired measurements per lab/subject, with median crosshairs.
  • — Drag a cutoff and watch sensitivity/specificity update live.

References & further reading

Other fields: · · · · · · · ·