A validation report is the document that records what was measured, what the measurements gave and what those results were judged against. The report is the evidence that a method was fit to put into service, and it is usually the only part of the work anyone reads again.
Three people will read it, and none of them watched the work being done. An inspector or assessor, checking that the work was done and the conclusion follows. A reviewer at a regulator, if the method is going into a submission. And a colleague, three years from now, who has inherited the method and needs to know why the acceptance limit was set where it was. A report that satisfies the third reader satisfies the other two.
The acceptance criterion belongs in the report before the results, and it belongs there because it was fixed before the study ran. A limit chosen after seeing the data is not a limit at all. A limit set afterwards only describes the data.
For each characteristic, state the limit, the source of the limit and the reasoning if the source needed interpretation. Analytical performance specifications come from the three Milan models: clinical outcome, biological variation and the state of the art. Regulatory and proficiency-testing limits can also bind, whatever the models say. The sources do not agree with each other. Which one was used is part of the finding.
Also fix the study design in advance: how many samples, how many days, how many replicates and at which concentrations. Sample size is not a detail. Sample size decides how wide the intervals come out. A report that does not say how the number was chosen invites the question at the worst moment.
Each performance characteristic gets its own section, and each section answers the same four questions: what was measured, what the estimate was, how uncertain that estimate is, and what limit it was judged against. The characteristics come from the framework you work under. The US Clinical Laboratory Improvement Amendments (CLIA) name four for a verification and seven for an establishment. ICH Q2(R2) names a different set for pharmaceutical analytical procedures.
| Characteristic | What the report states | Protocol |
|---|---|---|
| Precision | Repeatability and within-laboratory imprecision at each concentration, as SD and CV, with the degrees of freedom behind them | EP05 to establish, EP15 to verify |
| Trueness and bias | Bias at each medical decision point, each with a confidence interval, and the regression it was predicted from | EP09 |
| Method comparison | Which regression was fitted and why, slope and intercept with confidence intervals and a difference plot against the allowable difference | EP09 |
| Linearity | The deviation from linearity at each level, against the allowable nonlinearity band | EP06 |
| Measuring interval | The lower and upper limits, and which study set each one | EP06 with EP17 |
| Detection capability | LoB, LoD and LoQ, each with the error rate or imprecision goal that defines it | EP17 |
| Interference | Which interferents were tested, at which concentrations and the highest level at which the bias stayed inside the allowable band | EP07 |
| Reference interval | The limits with their confidence intervals, the method used, the number of reference individuals and the partitions | EP28 |
| Total analytical error | The estimate, and the allowable total error it was judged against | EP21 |
| Qualitative performance | Sensitivity and specificity against a reference standard, or positive and negative percent agreement against a comparator, with confidence intervals | EP12 |
A point estimate on its own tells a reader almost nothing about whether the method passed. A sensitivity of 90% from twenty diseased subjects and one from four hundred are the same number. The two are not the same finding. The first carries a confidence interval running down towards 70%.
So report the interval next to every estimate, and say which method produced it. Wilson for a proportion near the boundary, DeLong for an area under a ROC curve, a percentile bootstrap where an analytic interval does not exist. The method matters because intervals from different methods are different widths, and a reviewer who recomputes yours will get a different answer if they use a different one.
Judge the interval against the limit, not the point estimate. A bias of 2% against an allowable bias of 3% is a pass only if the interval also sits inside 3%. If the interval crosses the limit, the study cannot confirm a pass. Either the bias is too close to the limit or the study was too small to tell.
A comparison is only interpretable if the reader knows what the method was compared against. Name the comparative method, its manufacturer and its version, and say plainly whether it is a reference method or simply the assay currently in use. A reference method and the assay already in use are different claims. The second does not support the word accuracy in its metrological sense. CLIA’s “accuracy” is looser, and a comparison with the method in use is the usual way to verify it.
Describe the samples the same way: how many, patient or spiked or pooled, the matrix, the storage and how the concentrations were spread across the measuring interval. A comparison run entirely near the middle of the interval says nothing about the ends. Commutability belongs here too where processed materials were used.
Record the software and its version alongside the instrument and reagent lots. A number that cannot be recomputed is not evidence, and the calculation is part of the method.
Excluded data is the part of a report inspectors look for first. Account for every sample planned, measured, analysed and omitted. Every exclusion carries a reason, and the reason was decided by a rule rather than by looking at the result.
Record outlier handling as a rule applied, not as a judgement made: the test used, the criterion and how many observations it removed. Removing a point because it is inconvenient and removing it because a documented rule flagged it look identical in the final numbers and are not the same thing.
Where a study was repeated, the report says so, and says why. A first attempt that failed and a second that passed is a finding about the method. Reporting only the second hides that finding, and a reader who discovers it will doubt the rest of the report.
The worked examples behind each characteristic are listed in the guide for that characteristic. The roadmap of the EP protocols from the Clinical and Laboratory Standards Institute (CLSI) maps which study answers which question. Each guide it links carries the example workbook for its own analysis, ready to open in the Analyse-it trial.
An acceptance limit that appears after the results. A limit set once the data are in is unfalsifiable, and a reviewer will recognise it. Fix the criterion in the protocol, and if it has to change, record the change and the reason.
Point estimates without intervals. The most common gap of all. A table of slopes, biases and sensitivities with no confidence intervals gives the reader no way to tell a well-powered pass from an under-powered one.
Calling the comparator a reference method. If the comparative method is the assay you already run, the study measures agreement, not accuracy in the metrological sense. Under CLIA the same study is how accuracy is verified, so say which sense you mean.
Reporting r as evidence of agreement. A correlation coefficient near 1 is compatible with a large constant bias. Correlation is the wrong tool for method comparison, and its appearance in a validation report is usually a sign that no agreement analysis was done.
Silent exclusions. A sample count in the results that does not match the count in the design, with nothing explaining the difference, loses a reader’s confidence in everything else.
A conclusion that does not name the limit. “Performance was acceptable” is not a conclusion. “Bias at the 7.0 mmol/L decision point was 1.8%, interval 0.9% to 2.7%, against an allowable bias of 3.0%” is.
Analyse-it produces what the report needs, with the intervals attached, inside Excel:
Every feature from all five editions for 15 days. The analyses above sit in the Method Validation and Ultimate editions, from US$ 475 a year. Validated against NIST and CLSI reference datasets. See validation vs verification to settle how much of the above your method actually owes.