The correlation coefficient is still the most commonly reported number in a method comparison study, and almost always the wrong one. Open a validation report, a reagent evaluation, or a manuscript comparing a new assay against a predicate, and somewhere near the top you will find r. Usually with a reassuring string of nines after the decimal point, and usually offered as evidence that the two methods “agree well”.
A high correlation is evidence of no such thing. Bland and Altman made the point in 1986, and nearly forty years later the habit persists.
The correlation coefficient quantifies the strength of linear association between two variables: the tendency for one to increase as the other increases. Agreement is a different property entirely. It asks whether the two methods give the same value for the same sample.
The two diverge immediately. Suppose a new method reads exactly double the reference method across the whole measuring interval. Every point falls precisely on a straight line. The correlation is a perfect 1.00. Yet the methods disagree by 100% at every concentration. A constant offset behaves the same way. If the new method reads 5 units high everywhere, r is still 1.00 while a clinically important bias goes entirely undetected. Correlation is blind to the systematic errors (proportional and constant bias) that a method comparison exists to detect.
The point is easiest to see in one picture. Take any dataset where the new method reads consistently high, plot it, and compute r. The correlation is near-perfect. The difference plot of the same data shows the bias immediately. One dataset, two statistics, opposite conclusions. Only one of them is useful.
This is the more subtle problem. Range dependence is what makes r look like a property of the two methods, when it is really a property of your sample.
The correlation coefficient is, in effect, the ratio of the variation between your samples to the total variation, measurement error included. Spread your samples across a wider concentration range and the between-sample variation grows while the measurement error stays put. r rises, even though nothing about the methods’ agreement has changed. Pick a narrow range and the same methods will produce a mediocre-looking r.
The practical consequence is that a high correlation can be manufactured by selecting samples across a broad interval. Two studies of the same two methods can report very different coefficients purely because of sample selection. A number that moves around based on how you chose your specimens is not telling you about the methods.
Reporting a p-value alongside r only compounds the problem. The null hypothesis being tested is that the true correlation is zero: that the two methods are unrelated. Two procedures measuring the same analyte are, of course, related. Rejecting that null tells you nothing you did not already know before collecting a single sample. A tiny p-value here is not a finding. It merely restates the obvious.
There is something to this one, provided you are precise about what it claims. The heuristic is sometimes attributed to Stöckl and Thienpont, and is referenced in older CLSI guidance. It holds that once r exceeds about 0.975, the range of true values is wide enough relative to measurement error that the slope from ordinary least squares regression is not badly biased by error in the X variable.
Read carefully, that is a statement about range adequacy, not about agreement, and not a justification for using ordinary regression as your comparison. With a full complement of proper errors-in-variables procedures available, there is no reason to rely on a loose guideline about when an inappropriate model becomes tolerable. Use the coefficient, if at all, as a rough check that your samples span an adequate interval. Never as your principal result.
The alternatives are more informative and easier to interpret. They speak in the units your readers care about.
A difference plot, Bland and Altman’s, plots the difference between the methods against the best estimate of the true value. It shows the bias directly, whether that bias is constant or grows with concentration, and it exposes outliers and non-constant scatter that a correlation coefficient hides completely. No other single picture in a method comparison is as useful.
Limits of agreement turn that plot into an interval: the range within which most differences between the two methods are expected to fall. When one method is a reference, that interval is a usable estimate of total error.
An appropriate regression procedure estimates constant and proportional bias, and lets you predict that bias at the medical decision points that actually matter, with confidence intervals and equivalence tests against an allowable difference. Use Deming or Passing–Bablok rather than ordinary least squares, for reasons the regression guide covers.
None of these can be inflated by widening the sample range. None is blind to systematic error. Every one of them answers the real question: when these two methods measure the same sample, how far apart are the answers, and does that difference matter?
That is the question a correlation coefficient was never built to answer.
Analyse-it runs the analyses that answer the agreement question, on your own data, inside Excel:
Every feature from all five editions for 15 days. Method comparison is in the Method Validation and Ultimate editions, from US$ 475 a year. The Medical edition covers Bland–Altman agreement and bias at decision points, but not the regression fits, from US$ 340 a year. Validated against NIST and CLSI reference datasets. Full detail in the method comparison reference guide.