Outliers: When Exclusion Is Legitimate and When It Is Not
Look at any published figure and ask what you can tell about the numbers that are missing from it. The answer is essentially nothing. Points that were dropped leave no mark; removals often go unmentioned altogether; and where a paper does mention one, the three things worth knowing — how many went, on what grounds, and whether those grounds predated the data — are usually the three things absent. “One outlier was removed” records an action without defending it. A dataset presented whole, with an odd value flagged and unexplained, has served the reader better, even though the graph looks untidier for it.
That asymmetry is what makes deletion the costliest analytical decision in the set. Nobody downstream can audit it.
Two decisions that get the same name
Throwing out a run and throwing out a number are routinely spoken of as one practice. They are not.
Scrapping a whole run because a control failed is a verdict on the experiment, reached using evidence that lies outside the measurements under study. That version holds up. Deleting a single observation from a run that was otherwise fine is a verdict on that observation, and the evidence brought in support of it is, almost invariably, the observation itself. A rule composed in advance can govern the first case cleanly. It can seldom govern the second. Which is why criteria are worth writing at the level of the run, and exclusions at the level of the individual value are worth not making.
Independence is the whole of the test
A number may go when an assignable cause turns up that does not depend on the number: a pipetting error someone logged, a well with visible contamination, an instrument fault noted while it happened, a specimen known to have been mishandled.
Note the word independent, because it is carrying everything. “This point sits far from the rest” names no cause at all; it is merely what sent you searching. A cause unearthed only because a value looked wrong, and then accepted because it accounts for that value, is not independent evidence of anything. The same reasoning runs through an out-of-specification investigation, laid out in out-of-specification results and retesting, where the assignable cause must be located rather than presumed.
No amount of arithmetic can declare a number wrong
Various tests will flag extreme values. What they do is identify observations that would be improbable under an assumed distribution. That is the extent of it.
Two consequences. The assumption may not hold: analytical and biological data are often skewed or heavy in the tails, and something that looks improbable against a normal model can be perfectly commonplace under whatever distribution actually applies. And improbable does not mean mistaken. An unusual reading is sometimes the single most informative point in the set — evidence of a real subpopulation, a threshold being crossed, or genuine instability in the material. A test establishes that a value is unusual. Nothing in the calculation reaches the conclusion that it is erroneous.
What the deletion buys, and what it costs
Take out the most extreme point and the apparent spread shrinks. Error bars tighten. The p-value falls. Make a habit of it and every experiment in the record claims more precision than it earned.
Small sets make this acute. Drop one of three and two remain, which supports no meaningful estimate of variability whatever. Unfortunately that is exactly the situation where the pull toward deletion is strongest and the harm done is largest, for reasons set out in technical and biological replicates.
Four things to do instead
Run the analysis both ways and publish both. Conclusion holds either way, and the argument evaporates; conclusion flips, and that instability is itself the result. Alternatively, choose a summary that does not care much about extremes — a median with an interquartile range characterizes skewed data without anybody adjudicating what deserves to stay.
Collecting more data is the only route that adds information rather than subtracting it: one extreme value among twenty barely registers, whereas one among three runs the show. And treat the value as a lead worth following. An unexplained extreme in analytical work may be a contamination event of the sort covered in cross-contamination on the bench, which you want to know about whatever eventually becomes of the datum.
Criteria first, data second
The fix here is procedural, not statistical. Before anyone looks at a result, commit to paper what will invalidate a run:
- a suitability failure of the kind described in system suitability;
- a control falling outside a range fixed beforehand;
- a plate showing an edge effect;
- a culture at the wrong confluence.
Criteria written ahead of time can be applied without the outcome tugging at them. Criteria improvised afterwards cannot be, no matter how sensible they read — and sensibleness is no safeguard, because a plausible story can be assembled for virtually any individual number.
Where that ordering was not observed, two honest courses remain. Keep the value. Or state openly that it was removed after the fact, and on what reasoning. What is not honest is showing the trimmed set as though it were what came out of the experiment.
