Why correlation is the wrong statistic for method comparison The correlation coefficient is the most reported number in a method comparison study, and almost always the wrong one. Why r cannot tell you whether two methods agree, and what to report instead.

The correlation coefficient is still the most commonly reported number in a method comparison study, and almost always the wrong one. Open a validation report, a reagent evaluation, or a manuscript comparing a new assay against a predicate, and somewhere near the top you will find r. Usually with a reassuring string of nines after the decimal point, and usually offered as evidence that the two methods “agree well”.

A high correlation is evidence of no such thing. Bland and Altman made the point in 1986, and nearly forty years later the habit persists.

Correlation measures association, not agreement

The correlation coefficient quantifies the strength of linear association between two variables: the tendency for one to increase as the other increases. Agreement is a different property entirely. It asks whether the two methods give the same value for the same sample.

The two diverge immediately. Suppose a new method reads exactly double the reference method across the whole measuring interval. Every point falls precisely on a straight line. The correlation is a perfect 1.00. Yet the methods disagree by 100% at every concentration. A constant offset behaves the same way. If the new method reads 5 units high everywhere, r is still 1.00 while a clinically important bias goes entirely undetected. Correlation is blind to the systematic errors (proportional and constant bias) that a method comparison exists to detect.

The point is easiest to see in one picture. Take any dataset where the new method reads consistently high, plot it, and compute r. The correlation is near-perfect. The difference plot of the same data shows the bias immediately. One dataset, two statistics, opposite conclusions. Only one of them is useful.

Correlation depends on the range of samples you chose

This is the more subtle problem. Range dependence is what makes r look like a property of the two methods, when it is really a property of your sample.

The correlation coefficient is, in effect, the ratio of the variation between your samples to the total variation, measurement error included. Spread your samples across a wider concentration range and the between-sample variation grows while the measurement error stays put. r rises, even though nothing about the methods’ agreement has changed. Pick a narrow range and the same methods will produce a mediocre-looking r.

The practical consequence is that a high correlation can be manufactured by selecting samples across a broad interval. Two studies of the same two methods can report very different coefficients purely because of sample selection. A number that moves around based on how you chose your specimens is not telling you about the methods.

Two panels with the same underlying relationship and scatter, sampled over a narrow range (correlation 0.76) and a wide range (correlation 0.97).
The same method relationship and the same measurement scatter, sampled over a narrow range and a wide one. The correlation coefficient climbs from 0.76 to 0.97. That is a sign that r reflects your sample’s range as much as the methods' agreement.

The significance test is meaningless here

Reporting a p-value alongside r only compounds the problem. The null hypothesis being tested is that the true correlation is zero: that the two methods are unrelated. Two procedures measuring the same analyte are, of course, related. Rejecting that null tells you nothing you did not already know before collecting a single sample. A tiny p-value here is not a finding. It merely restates the obvious.

“But r was above 0.975, so ordinary regression is fine”

There is something to this one, provided you are precise about what it claims. The heuristic is sometimes attributed to Stöckl and Thienpont, and is referenced in older CLSI guidance. It holds that once r exceeds about 0.975, the range of true values is wide enough relative to measurement error that the slope from ordinary least squares regression is not badly biased by error in the X variable.

Read carefully, that is a statement about range adequacy, not about agreement, and not a justification for using ordinary regression as your comparison. With a full complement of proper errors-in-variables procedures available, there is no reason to rely on a loose guideline about when an inappropriate model becomes tolerable. Use the coefficient, if at all, as a rough check that your samples span an adequate interval. Never as your principal result.

What to report instead

The alternatives are more informative and easier to interpret. They speak in the units your readers care about.

A difference plot, Bland and Altman’s, plots the difference between the methods against the best estimate of the true value. It shows the bias directly, whether that bias is constant or grows with concentration, and it exposes outliers and non-constant scatter that a correlation coefficient hides completely. No other single picture in a method comparison is as useful.

The same method-comparison data shown two ways: a scatter plot of one method against the other, and a Bland-Altman difference plot of the differences against their mean with the mean bias and allowable-difference bands.
The same data shown two ways (CLSI EP09-A3, Example 1): a scatter plot, and the Bland–Altman difference plot beside it. The scatter looks like a reassuring straight line. The difference plot exposes the bias a correlation coefficient hides.

Limits of agreement turn that plot into an interval: the range within which most differences between the two methods are expected to fall. When one method is a reference, that interval is a usable estimate of total error.

An appropriate regression procedure estimates constant and proportional bias, and lets you predict that bias at the medical decision points that actually matter, with confidence intervals and equivalence tests against an allowable difference. Use Deming or Passing–Bablok rather than ordinary least squares, for reasons the regression guide covers.

None of these can be inflated by widening the sample range. None is blind to systematic error. Every one of them answers the real question: when these two methods measure the same sample, how far apart are the answers, and does that difference matter?

That is the question a correlation coefficient was never built to answer.

Do it properly with Analyse-it

Analyse-it runs the analyses that answer the agreement question, on your own data, inside Excel:

Every feature from all five editions for 15 days. Method comparison is in the Method Validation and Ultimate editions, from US$ 475 a year. The Medical edition covers Bland–Altman agreement and bias at decision points, but not the regression fits, from US$ 340 a year. Validated against NIST and CLSI reference datasets. Full detail in the method comparison reference guide.