Bland–Altman explained: limits of agreement The difference plot answers the question a method comparison is really asking — when these two methods measure the same sample, how far apart are the answers, and is that gap acceptable? How to build it, read it, and know its limits.

Correlation cannot tell you whether two methods agree. Regression characterises the bias but not, at a glance, its size in the units you care about. The Bland–Altman difference plot answers the question a method comparison exists to ask: how far apart the two methods’ answers are, across the measuring range, in the analyte’s own units. The plot is the most useful single picture in a method comparison, and also the one most often read without care.

How the difference plot is built

For each sample, plot the difference between the two methods on the vertical axis against the mean of the two on the horizontal. The mean is used as the best available estimate of the true value when neither method is a reference. Plotting the difference against one of the methods instead would build in a spurious relationship. When one method is a true reference, plotting the difference against that reference is legitimate and often preferred.

That single change of axes makes bias visible. A scatter plot of one method against the other hides a constant offset inside a cloud near the line of equality. The difference plot puts the offset on the vertical axis where you can see it.

Mean bias and the limits of agreement

Two numbers summarise the plot. The mean difference is the average bias between the methods, and a value away from zero means one method reads systematically higher than the other. The limits of agreement are that mean difference plus and minus 1.96 standard deviations of the differences:

limits of agreement = mean difference ± 1.96 × SD of the differences

The multiplier is 1.96 because 95% of a normal distribution lies within 1.96 standard deviations of its mean. The two limits therefore bracket about 95% of the differences between the methods. Note that the standard deviation is the one of the differences, not of either method’s own results.

The limits are what you report, because they say in plain units how far apart the two methods are likely to be for an individual sample. Compare them against a clinically allowable difference. Limits that fall inside what is clinically tolerable mean the methods can be used interchangeably. Limits that fall outside it mean they cannot, however small the mean bias, because a near-zero average difference counts for little when individual samples still disagree by more than matters.

A Bland-Altman difference plot schematic: differences plotted against the mean of the two methods, a solid mean-bias line, and two dashed limit-of-agreement lines labelled mean plus 1.96 SD and mean minus 1.96 SD, with a dashed zero reference line.
How the limits are built. The mean-bias line is the average difference; the limits of agreement sit 1.96 standard deviations either side of it, so about 95% of differences fall between them. Whether that spread is acceptable is judged against the clinically allowable difference, not against zero.

Check the assumptions the limits rely on

Constant limits of agreement assume two things, and the plot itself lets you check both. The differences should be roughly normally distributed, and the bias and the scatter should be constant across the measuring range. A violation shows up immediately in the picture. Where the cloud of differences fans out as concentration rises, so that disagreements are larger at high values, constant limits are too wide at the bottom of the range and too narrow at the top. Quoting a single pair of limits then misleads the reader.

Handling proportional bias and non-constant scatter

When the scatter grows with concentration, you have two options. Transform the data. A log transformation often stabilises the spread, after which constant limits on the log scale become proportional limits back on the original scale.

Or model the relationship. Regress the differences on the mean and let the limits of agreement widen with concentration, following the data rather than pretending the spread is flat. A percentage difference plot, showing the difference as a proportion of the mean, is a third view that often makes a proportional pattern obvious. Do not fit constant limits to a fanning plot and report them as if they held everywhere.

The limits are estimates too

The limits of agreement are computed from a sample, so they carry their own uncertainty — more than most people expect. The confidence interval around a limit is wide unless the sample is substantial, which is why a method comparison needs a reasonable number of samples. The mean bias settles down quickly as samples accumulate. The limits need considerably more data before they can be estimated tightly enough to rely on. Report the confidence intervals alongside the limits.

What to report

A Bland–Altman analysis is under-reported more often than it is done badly. Bland and Altman’s own papers, and the reporting guidance that follows them, ask for a specific set of items. Report all of these:

  • The number of samples, and how they were chosen — the range they span decides what the limits describe.
  • The mean difference with its confidence interval, in the analyte’s units, stating which method was subtracted from which.
  • Both limits of agreement, each with its own confidence interval. The limits are estimates, and their intervals are wide at realistic sample sizes.
  • The allowable difference you judged against, and where it came from — a clinical requirement, an allowable total error, or a specification derived from biological variation.
  • The plot itself, with the mean-bias line, both limits, and the allowable difference band drawn on it.
  • Whether the assumptions held: were the differences roughly normal, and were the bias and scatter constant across the range? If not, say what you did instead.
  • Any replicate structure. Measuring each sample once on each method is not the same analysis as measuring it several times, and the standard formula assumes the former.

Two omissions account for most weak reports. The first is quoting limits of agreement with no allowable difference to judge them against, which leaves the reader unable to tell whether the methods agree well enough. The second is quoting the limits without their confidence intervals, which presents an uncertain estimate as though it were exact.

Limits of agreement are not confidence intervals

The two get confused, and they describe different things.

The limits of agreement describe the spread of individual differences: how far apart the two methods are likely to be for one sample. They stay roughly the same width however many samples you collect, because they describe the methods, not your study.

A confidence interval describes uncertainty in an estimate. There is one around the mean difference, and one around each limit of agreement, and both narrow as the sample grows. A confidence interval on a limit answers “where might this limit really sit?”, while the limit itself answers “how far apart might two measurements of one sample be?”.

So a study can have a tight confidence interval around a limit of agreement that is far too wide to accept, or a tolerable limit estimated so imprecisely that you cannot rely on it. Report both, and judge the limit against the allowable difference while judging the interval against how much precision your decision needs.

Bland–Altman or regression?

The two answer different questions and work best together. Bland–Altman asks “how far apart are the methods, and is that acceptable?”, judging agreement in clinical units against an allowable difference. Regression — Deming or Passing–Bablok — asks “what is the structure of the bias?”, separating a constant offset from a proportional one and estimating the bias at specific decision points. A thorough comparison runs both. Use the difference plot to judge whether the disagreement is tolerable. Use the regression to characterise where it comes from.

Downloads

Download the CLSI EP21-A Bland–Altman example workbook (.xlsx) — a sodium agreement analysis with the mean difference, limits of agreement, and an allowable difference band, ready to open in the Analyse-it trial.

A difference plot of sodium measured by a candidate method against the average of the two methods, with the mean difference line, 95% limits of agreement at minus 2.942 to 2.777, and an allowable difference band of plus or minus 4 mmol per litre; below it a table giving the mean difference as minus 0.083 with its confidence interval.
A difference plot with limits of agreement (CLSI EP21-A, sodium). The difference is plotted against the average of the two methods, as it should be. The mean difference is −0.083 mmol/L and the limits sit at −2.942 and 2.777 — comfortably inside the allowable difference of ±4 mmol/L, so the two methods agree closely enough to be used interchangeably.

Common mistakes

Reporting limits of agreement with nothing to judge them against. Limits of −2.9 to 2.8 mmol/L are neither good nor bad until you state the allowable difference. Without it the reader cannot reach a conclusion.

Plotting the difference against one method instead of the mean. Unless that method is a true reference, this builds in a spurious slope. Use the mean of the two.

Reporting only the mean bias. A small average difference can hide wide limits of agreement. The limits, not the mean, are what you compare against the allowable difference.

Fitting constant limits to a fanning plot. If the scatter grows with concentration, constant limits are wrong across most of the range. Transform, or let the limits vary with the mean.

Using it unchanged to compare two instruments. The plot is right, but for two analysers running the same method the question is whether each difference sits inside a pre-set limit rather than what the limits of agreement are. See comparing instruments within one laboratory.

Quoting the limits without their confidence intervals. The limits of agreement are uncertain, especially with modest samples. Report the intervals, and size the study to make them tight.

Draw a Bland–Altman plot with Analyse-it

Analyse-it builds the difference plot and the limits from your own pairs, inside Excel:

  • Mean bias and 95% limits of agreement, each with its own confidence interval
  • Percentage difference plots, and proportional limits where the scatter is not constant
  • An allowable difference band drawn on the plot, so the judgement is visible rather than inferred

Every feature from all five editions for 15 days. Method agreement is in the Medical, Method Validation and Ultimate editions, from US$ 340 a year. Validated against NIST and CLSI reference datasets. Full detail in the method comparison reference guide; see also why correlation fails and choosing a regression.