Error Bars: Standard Deviation, Standard Error and Confidence Intervals
Three wells were filled from one dilution of one preparation. The figure reports n equal to three. Every error bar on that chart is now roughly a third shorter than anything the experiment could honestly justify, and no amount of care about which statistic was plotted will repair it. Most misleading error bars are born here, in the counting, rather than in the choice of formula.
The count decides everything downstream
Both of the quantities researchers commonly plot are computed from a sample size, and that sample size turns entirely on what qualifies as an independent observation. Three wells drawn off a single dilution are three looks at the same thing, not three things. The distinction is worked through in technical and biological replicates, and it is upstream of every other problem discussed below.
Three quantities, three questions
A caption naming the bars resolves the ambiguity; a caption that omits it leaves a reader guessing between quantities that can differ by a factor of two or more.
| Quantity | What it describes |
| Standard deviation | How widely the individual observations scatter. Collect more data and this does not contract; the estimate of it sharpens, but the scatter is whatever the system produces |
| Standard error of the mean | How well the average has been pinned down. Divide the standard deviation by the square root of the sample size, so it contracts as the sample grows even when nothing about the underlying variability has moved |
| Confidence interval | A range built so that, across repeated experiments, some stated fraction of such ranges would enclose the true value. For a large sample it spans about two standard errors on each side of the mean, and it widens when the sample is small |
One of these characterizes how variable the system is. The other two characterize how confidently the average is known.
Why the shortest bar tends to win
Standard error is the narrowest option available, undercutting standard deviation by the square root of the sample size, so the identical dataset simply looks firmer when drawn that way.
Choosing it is not automatically a fault. Where the claim being made is about a mean, the precision of that mean is genuinely the quantity at issue. The trouble starts when nothing on the figure says which was plotted, since a reader trained on one convention will silently apply it to the other. Four replicates and an unlabeled bar leave two candidate quantities in play, differing twofold.
Overlap does not mean what it appears to mean
Two intuitions circulate, and neither survives inspection. Standard error bars that fail to touch do not demonstrate a significant difference. Standard error bars that do touch do not exclude one. How overlap relates to any test depends on the sample sizes involved and on which test was run.
Intervals do somewhat better here: a pair of ninety-five percent intervals sitting clear of each other generally corresponds to a difference that would test as significant. Even so, the dependable move is to compute the interval around the difference between groups, since the difference is what the claim was about in the first place.
What three observations can support
At three or four data points, the standard deviation is itself a very rough estimate, which makes any standard error derived from it unreliable and any confidence interval built on top of that both wide and unstable.
Saying so plainly matters, because n equal to three is nearly the default across this literature. Bars computed from three numbers have the appearance of precision without the substance; what they really convey is that somebody ran the measurement more than once. The same limitation shadows the does-this-sample-represent-the-batch problem in sampling plans, where a small sample bounds the available claims no matter how the result is drawn.
The caption that almost never appears
Four items belong there. Which quantity the bars show. What n equals, together with what one observation consisted of. Whether those observations were independent of each other. And whether the plotted center is a mean or a median, which starts to matter the moment the data skew, as concentration-response data routinely do.
When replicates are few, abandoning the bar and plotting the points themselves is both more honest and no larger on the page. Readers then see the actual scatter rather than reconstructing it from a summary that may have been selected for how it looked.
The same failure on an analytical document
Certificates are governed by the identical principle. State a purity figure with no uncertainty attached and a point estimate has been dressed up as an exact value, which is the theme of measurement uncertainty. Carry that figure to more decimals than the method can sustain and a precision has been asserted that does not exist, as significant figures and rounding sets out.
Chart and certificate fail the same way: a presentation decision makes a number appear better determined than the underlying work allows. One fix serves both. Say what the spread stands for, and say how many observations went into producing it.
