Out-of-Specification Results and What a Retest Can Prove
Hold a certificate up and ask it how many times the assay was run. It will not tell you. Nor will it say whether any determination was set aside, or what an investigation decided, or whether an investigation took place. All it carries is the figure that was reported at the end. A spotless certificate is equally consistent with a clean first measurement and with weeks of documented inquiry that eventually justified invalidating an earlier one.
No scandal in that; reporting a lot against its limits is the whole job of the document. The consequence is simply that most of the confidence a certificate carries rests on whether the laboratory behind it operates a defined procedure for handling failures at all. Accreditation is the usual outside check on that question, subject to the scope caveats in what ISO/IEC 17025 accreditation covers, and the incentives differ by arrangement in the way described in third-party versus in-house testing.
Data does not become un-data
Here is the rule everything else follows from. The moment an instrument produces a number, that number is data. A later measurement, however much more convenient, cannot retract it. A laboratory that keeps testing and keeps discarding until something passes has stopped measuring and started choosing.
This is not pedantry about recordkeeping. Take a lot that genuinely sits on its limit: run the assay often enough and chance alone will eventually hand you a value on the right side of the line. Publish that value alone and a coin toss has been laundered into a certificate.
The order an investigation has to follow
When a result fails, the formal answer is an investigation, and its object is to work out whether what failed was the material or the measurement. The order is not arbitrary. Everything begins in the laboratory: was there a traceable mistake — a standard weighed wrong, a mobile phase made up incorrectly, a suitability sequence that did not pass, a number copied across badly? Those checks are the ordinary ones set out in system suitability.
Whatever turns up there has to be found, not supposed. Saying the instrument probably drifted identifies nothing; a calibration record that shows the drift identifies something. Only once the laboratory phase closes without a cause does attention move to the material. At that point the result is a property of the sample, and the live question becomes whether the sample stood in fairly for the lot — which is the subject of sampling plans and batch representation.
One legitimate use of a retest, three impostors
Retesting has a narrow proper role, and three misuses that are easy to mistake for it. It is not an appeal: a second number does not overrule the first. Where a laboratory error was identified and fixed, the repeat stands in for a result that has been invalidated; where nothing was found, both numbers are data and the material is out of specification on the weight of them.
It is not a search either. Running the assay until a passing figure surfaces — testing into compliance, as it is sometimes called — turns the logic of measurement inside out. And it is no substitute for sampling more widely: reinjecting one vial interrogates the injection, re-preparing from the same aliquot interrogates the preparation, and only fresh material drawn from the lot interrogates the lot.
The number of retests allowed, where the material for them comes from, and whose signature authorizes them all belong in the written procedure before anybody sees a result. A retest allowance settled on after the first failure is not a rule. It is a negotiation.
The 97.2 and 99.0 problem
Suppose two determinations come in at 97.2% and 99.0%, and the pair is averaged to 98.1%. The lot is now reported as compliant, and a spread of nearly two percent — in the method, the sample or the material, and it matters which — has disappeared from view. That spread is far larger than the uncertainty discussed in measurement uncertainty, and it was the most interesting thing the testing found.
Averaging is fine where it was designed into the method from the outset. A method that calls for the mean of three preparations reports that mean, and no single preparation is a result on its own. What is not fine is a method calling for one determination, a failure, a second run, and a quiet average of the two.
Three ways a number can be wrong; one breaches a limit
| Category | What happened | Example |
| Out of specification | Falls outside an acceptance criterion | Purity below the stated limit |
| Out of trend | Passes, but departs from lot history | 98.2% where history clusters at 99.4% |
| Out of expectation | Anomalous with no limit attached | Unexpected peak, shifted retention time, unpredicted mass |
The middle row still signals that something changed, in the manner covered in lot-to-lot variation and trending. Neither the middle nor the bottom row will ever reach a certificate, since a certificate weighs one lot against its own limits. Both are visible only to whoever keeps the history, and that is one of the structural blind spots in reading any single report.
Three honest endings
When a lot has genuinely failed, the menu is short, and reporting a nicer figure is not on it:
- Reject the lot.
- Reprocess it, then test it as a new lot under a new identifier, per batch and lot numbering.
- Release it against a different, lower specification, with that specification printed on the report.
The third is perfectly proper and appears far less often than it ought to. Material called 94% against a 94% limit gives a buyer something to work with. The same material called 98% because the third determination obliged does not.
