What a p-Value Does and Does Not Tell You
Line up twenty comparisons in which nothing whatsoever is going on, apply the customary 0.05 cutoff, and on average one of them will come back labeled significant. That is arithmetic. Nobody cheated. The trouble starts when all twenty get run, one gets written up, and the reader is handed a verdict with no way of knowing how many attempts stood behind it.
Concentrations, time points, readouts — a study that varies all three is conducting a great many comparisons regardless of whether the write-up frames them that way. How many tests took place is information the reader needs and rarely receives. Formal corrections exist, but they are blunt and they cost power. Nominating the primary comparison before the data arrive, and labeling everything else exploratory, does more good: it is also what makes a finding worth an attempt at replication, in the sense of why two laboratories get different results.
Small studies exaggerate, even when everyone is honest
There is a sharper version of this, and it deserves naming. Where a study lacks power, the only findings capable of clearing a significance threshold are those in which the observed effect landed, by chance, on the large side. Modest true effects in underpowered work simply fail to reach the line and never appear. What accumulates in the record is therefore biased upward in magnitude, with no dishonesty required anywhere.
Peptide work sits badly exposed to this. The studies run small. They carry multiple readouts. They appear one at a time rather than as coordinated programs. Every one of those characteristics lifts the share of published positives that later fail to hold. It is a good part of why the reading discipline in what the published literature actually shows is necessary rather than fussy.
Back to the definition, which is narrower than the usage
Formally: the probability of data at least as extreme as the data in hand, on the assumption that the null hypothesis holds. Everything hinges on that conditional clause.
What it gives you is a statement about data, conditioned on a hypothesis. What people take from it is a statement about a hypothesis, conditioned on data. These are not the same quantity and one does not convert into the other. So 0.03 is not a three percent chance of a fluke, and it is not ninety-seven percent confidence that something real is present.
Four readings the quantity will not bear
The first treats the figure as the probability the hypothesis is false. Impossible by construction, since the calculation began by assuming the null to be true. The second promotes 0.05 into a boundary separating real effects from imaginary ones; it is a convention, picked without principle and kept out of habit, and nothing in nature distinguishes 0.049 from 0.051.
The third reads a non-significant outcome as demonstrating absence. But a null result looks identical whether the effect is genuinely zero or the study was simply too small to resolve it, and separating those two requires an interval, which a p-value is not. The fourth takes a smaller figure as evidence of a larger effect. Magnitude is not what it measures: a negligible difference pinned down with great precision yields a very small p-value, while a substantial difference measured sloppily may yield nothing at all.
The pair of numbers that would have answered the question
What anyone actually wants to know is how big the difference was and how well it was pinned down. Neither is available from a p-value. An effect size carrying a confidence interval supplies both together, and it fails gracefully as well — an interval that is wide and straddles zero announces that the study settled nothing, which tells a reader considerably more than the phrase “not significant” does. The parallel with reporting an analytical figure alongside its uncertainty, rather than stripped bare, is exact; see measurement uncertainty.
What to check before believing a result
- Is an effect size given anywhere, or does the paper report only significance?
- Does an interval travel with that effect size?
- How many comparisons were performed in total — not how many were shown?
- Was the sample size fixed before data collection began?
- Is the analysis presented the one that was planned, or one settled on once the data were in view?
A report offering “p < 0.05” with no magnitude attached and no denominator disclosed has delivered a judgment while withholding the evidence. Structurally it is the same failing as a certificate that prints “complies” where a measured value belongs, a point developed in how a specification limit gets set.
Where that leaves the statistic
Here is the fair verdict. This is a thin, conditional, narrowly scoped quantity that has been made to bear the weight of scientific conclusions for the better part of a century. It is not worthless: it does register that a pattern would be surprising under one particular null model. But it is the least informative of the several inputs that ought to go into a judgment.
In practice, then: go to the magnitudes and the intervals first, let the p-value sit in the footnotes where it belongs, and read one significant finding from one small study as an invitation to investigate further rather than as a matter settled.
