Technical and Biological Replicates Are Not Interchangeable
Few phrases in a methods section carry less information than “experiments were performed in triplicate.” Three wells drawn from one tube gets written that way. So does three experiments run on three separate days. The statistics that follow mean entirely different things in the two cases, and the reporting almost never tells you which one you are looking at.
Ask what one observation was
That single question resolves most of the confusion. Everything downstream depends on where the data point boundary falls.
Repeated measurement of one biological sample gives you a technical replicate. Three wells filled from a single dilution qualify, as do three injections drawn from one vial and three reads of the same plate. Such replicates characterize how precise the measurement procedure is.
An independent instance of whatever is under study gives you a biological replicate. Separate cultures, separately treated, preferably on different days and from different passages. These characterize how variable the system is.
Neither is superior; they answer unrelated questions. Trouble starts only when technical replicates get tallied up as the sample size supporting a statement about biology. So when reading, look for whether replicates happened on separate days, whether separate cultures or passages went into them, whether the compound was diluted afresh for each, and whether the stated n counts wells or counts experiments. Papers that specify all four can be assessed. Papers that specify none have left reproducibility unaddressed, which is the wider issue examined in why two laboratories get different results.
The error already has a name
Treating non-independent observations as independent is called pseudoreplication in the statistical literature. Ecologists described it decades back, where the identical blunder took the form of counting repeated measurements from one site as though they were replicate sites. The mistake does not change shape when it moves fields.
Having a name for it helps, because the name points straight at the cure. No alternative test fixes pseudoreplication and no correction factor patches it. The analysis simply has to happen at whatever level the claim addresses. A claim about cultures gets analyzed across cultures, with measurements inside each culture averaged beforehand. When both levels really do carry information, the appropriate instrument is a model that represents the nesting explicitly, and averaging is the crude approximation to that model.
Why the arithmetic flatters you
By construction, technical replicates resemble one another more than biological replicates do. They came out of the same preparation, the same dilution, the same pair of hands, the same few minutes.
Promoting them to independent observations therefore does two things simultaneously. Variability gets understated, since whatever sources of variation the group shares cannot appear within the group. Sample size gets overstated. Both effects drive the standard error downward, which shrinks the error bars and makes any test likelier to cross its threshold, exactly as described in error bars.
Out comes a figure that appears precisely determined next to a p-value that appears persuasive. Neither says anything about how reliably the finding would repeat.
A gradient, not a binary
The real test is whether replicates hold some source of variation in common that the claim was meant to average over. A few arrangements are unambiguous; others sit in between.
| Three wells, one dilution, one plate | Technical, with no room for argument |
| Three plates, one dilution, one day | Technical as far as the preparation goes, since dilution error and any adsorptive loss of the type covered in adsorptive loss to surfaces are shared by all three |
| Three cultures, three dilutions, three days | Biological |
| Three separate cultures, all dosed from one stock dilution made once | Partly shared. Biology independent, compound preparation not. A bad preparation makes all three wrong at once. |
That final arrangement is both the most common and the most uncomfortable. It sits above a technical triplicate and below a genuine biological one, and the honest way to report it is to state plainly what the replicates had in common.
Different jobs for different replicates
Average the technical replicates. Their mean turns into a single observation, and their scatter serves as a quality check on the measurement rather than as data describing the system. Substantial disagreement among technical replicates signals an assay fault, whether in pipetting, per pipetting accuracy, or in plate position effects. Fix it; do not average it into invisibility.
Biological replicates are the observations themselves. How many there are is the sample size, and statistics get computed across them. Compressed to a sentence: technical replicates sharpen each point, biological replicates are the things you count.
The trouble with three
Three is a convention, not a derivation. It happens to be the fewest replicates from which a variance can be estimated at all, and the estimate it yields is poor.
Low power follows, and low power means the effects an experiment picks up are disproportionately those that happened to look big, which is the selection problem laid out in what a p-value does and does not tell you. Adding biological replicates helps with this. Adding technical replicates does not, because they contribute nothing about the variability that constrains the conclusion.
The same logic on an instrument
None of this is peculiar to biology. Inject one vial three times and you have measured the instrument. Prepare three solutions from a single weighing and you have measured the preparation. Take three samples from a batch and you have measured the batch, which is the distinction underlying sampling plans and batch representation.
The discipline is the same on either side. Establish what the claim is about, then count only those replicates that vary in the way the claim requires.
