Method validation follows a well-trodden path defined by the CLSI EP (Evaluation Protocols) series — establish how precise the assay is (EP05), whether it reads true and meets its claims (EP15), over what range it stays linear (EP06), how low it can reliably detect (EP17), what a normal result looks like (EP28), and whether it agrees with the method it replaces (EP09). This guide maps each step to the right tool, names its CLSI protocol and experimental design, and explains how to judge the result.
Two ideas run through all of it. First, the design matters as much as the statistic — a precision study needs replicates spread across days and runs, a comparison needs ~40 patient samples spanning the reportable range. Second, "pass or fail" needs a pre-defined acceptance limit: an analytical performance specification (APS), ideally derived from biological variation or clinical need rather than picked ad hoc.
Tip: these tools expect your data laid out as columns (measurements, run/day labels, comparison-method values), in the assay’s own units.
Run replicates across days and runs, then decompose the scatter into repeatability (within-run), between-run, between-day, and overall within-laboratory imprecision — each as an SD and CV% with confidence intervals, plus a measurements-by-day plot.
The full EP05-A3 precision experiment uses a nested 20-day × 2-run × 2-replicate design (80 measurements) at two or more concentrations. To verify a manufacturer’s precision claim more cheaply, the EP15-A3/ED3 protocol runs 5 replicates on each of 5 days (the "5×5" design). Judge the observed CV against the claim, or against a desirable imprecision goal (biological-variation APS: desirable CV ≈ 0.5 × within-subject variation CVI).
To decide whether a new assay can replace an existing one, measure ~40 patient samples that span the reportable range on both methods — ideally in duplicate over several days — and regress them against each other. Passing-Bablok is the robust, distribution-free default; Deming (and weighted Deming for proportional error) handles error in both axes; Bland-Altman shows the bias directly.
Don’t stop at "slope ≈ 1, intercept ≈ 0". The clinically relevant question is the predicted bias at the medical decision points: estimate it from the fit and compare against your allowable bias. Surveys of real labs find this "critical difference" step is the one most often skipped — yet it is what actually decides fitness for use.
Run a dilution/admixture series of ≥5 levels spanning the claimed range (in replicate) and check whether recovery stays on a straight line. The current EP06-ED2 approach fits first-, second-, and third-order polynomials and tests whether the non-linear terms matter; the tool reports the deviation from linearity at each level.
Statistical significance alone is not the verdict — judge each level’s allowable deviation from linearity (ADL) against a limit, commonly ≤ 0.5 × TEa or the allowable bias. If you rely on the manufacturer’s dilution factor, verify it rather than assuming it (a common shortcut worth avoiding).
From blank and low-level replicates, estimate the Limit of Blank, Limit of Detection, and Limit of Quantitation. Following EP17-A2, LoB is the 95th percentile of the blank measurements (parametric: mean + 1.645·SD, from ≥60 blanks); LoD adds the low-level material’s own scatter; LoQ is the lowest concentration meeting a target CV or total-error goal.
Establishing an interval de novo needs ≥120 reference individuals; EP28-A3 then takes the central 95% by the non-parametric (N+1)p percentile with rank-based confidence limits (parametric, robust, and Box-Cox variants help with small or skewed samples). More practically, adopt a published interval and verify it with ~20 samples — it passes if not significantly more than the expected ~5% fall outside (an exact binomial test).
Round out validation and keep the assay honest in routine use. Screen a suspected interferent (EP07) with a paired test-vs-control design against an allowable bias. Combine bias and imprecision into a total analytical error and compare it to the allowable TEa (EP21), summarized as a sigma metric that also sets how much QC you need — high sigma tolerates simple rules, low sigma demands multi-rule QC. Day to day, Levey-Jennings / Westgard control charts and capability indices flag drift.
Every step above needs a pre-defined pass/fail limit — an analytical performance specification (APS). The 2014 Milan consensus ranks three ways to set one, best first: (1) the effect on clinical outcomes; (2) biological variation; (3) state of the art (what good labs currently achieve). Most quantitative analytes use the biological-variation model.
From the within-subject (CVI) and between-subject (CVG) biological variation — look them up in the EFLM Biological Variation Database — the desirable goals are: imprecision ≤ 0.5·CVI, bias ≤ 0.25·√(CVI²+CVG²), and total error TEa ≤ 1.65·(0.5·CVI) + 0.25·√(CVI²+CVG²). Feed these into the precision, bias, linearity (ADL ≤ 0.5·TEa), and total-error checks so that "acceptable" means something objective rather than eyeballed.