Robustness testing A method that works when its developer runs it is not the same as a method that works. Robustness asks what happens when the incubation runs long, the pH drifts, and a different analyst uses a different lot — all at once.

Robustness is the deliberate opposite of careful work. You vary the method’s parameters on purpose, by small amounts representing the drift that will happen anyway in routine use. Then you check that the result does not move more than you can tolerate. Temperature, pH, flow rate, wavelength, incubation time, reagent lot, column batch, analyst: whichever of these the method is plausibly sensitive to.

Robustness is a validation characteristic in its own right, and it is where system suitability criteria come from. A parameter the method turns out to be sensitive to is one you have to control tightly and say so in the procedure. A parameter it tolerates is one you can stop worrying about.

One factor at a time answers the wrong question

The intuitive approach is to hold everything fixed, move one parameter, see what happens, put it back, and move the next. One factor at a time is inefficient, and worse, it cannot detect an interaction.

An interaction is where the effect of one factor depends on the level of another. The method tolerates a slightly high temperature. It tolerates a slightly low pH. But at high temperature and low pH together the recovery collapses. Varying one at a time, each looks harmless. The combination that actually breaks the method is never tested. Routine use varies several parameters simultaneously, so that combination is exactly what will happen eventually.

Two interaction plots of recovery against incubation temperature at two pH levels. On the left the two lines stay parallel, so each factor can be judged on its own. On the right they diverge: recovery holds at pH 7.4 but collapses at pH 7.0 when the temperature rises, so the effect of temperature depends on pH.
The signature of an interaction is non-parallel lines. Vary temperature alone and it looks harmless. Vary pH alone and it looks harmless. Only varying them together exposes the combination that breaks the method.

Vary them together

A factorial design tests every combination of the factor levels. With three factors at two levels each (low and high, either side of nominal), that is eight runs. Those eight runs estimate all three main effects and every interaction between them. Compared with one-factor-at-a-time, the factorial design gives more information from a similar number of runs. Every run contributes to the estimate of every effect rather than to just one.

Where there are many candidate parameters, a screening design tests a fraction of the combinations to identify which few factors matter. Some higher-order interactions cannot be separated, which is usually the right trade, because robustness is a screening question. The goal is to find the parameters that need controlling, not to characterise the response surface in detail.

Choose the levels to represent realistic deviation, not extremes. If the procedure says 37 °C and the incubator holds ±1 °C, test 36 and 38. Testing 30 and 45 will certainly find an effect, and will tell you nothing about whether the method is robust in service.

The analysis is a multi-factor model

Fit a model with each factor as a term and the crossed terms for their interactions. The model gives the effect of each factor on the result, the effect of each combination, and a test of each against the residual variation. A main effects plot shows the size and direction of each factor’s influence. An interaction plot shows where two factors are not independent. Non-parallel lines are the signature.

Read the effect sizes before the p-values. With enough replication a trivial effect becomes statistically significant. A significant effect that shifts the result by a tenth of your allowable variation is not a robustness problem. The question is not whether a factor has any effect, since everything has some effect. The question is whether that effect, over the realistic range, is large enough to matter. Judge it against a limit you set beforehand, from an allowable total error figure or an equivalent specification.

Conversely, a large effect that misses significance because the study was small is not a clean bill of health. Report the effect with its confidence interval and judge that interval against your limit, rather than treating a p-value as the answer.

An effect of terms table listing three factors A, B and C with every two-way and three-way interaction, each with its sum of squares, F statistic and p-value, followed by a grid of two-way interaction plots showing the response against each factor at both levels of another.
A three-factor design analysed in one model (surface finish, Montgomery). Every main effect and every interaction is estimated from the same sixteen runs. Factor A dominates at F = 18.69; B follows at 4.33 and the A×B interaction at 3.10, neither reaching significance in a study this size — which is why the effect sizes matter more than the p-values. The interaction plots below show the same information as shape: parallel lines mean the factors act independently, converging or crossing lines mean they do not.

What to do with a factor that matters

A factor with an effect too large to tolerate is not a failed method. It is a parameter that needs a tighter specification, and the robustness study has just told you how tight. Narrow the permitted range until the effect across it falls inside your limit. Then write that range into the procedure and, where it can be monitored, into the system suitability criteria.

The factors that turn out not to matter are equally useful. They justify the tolerances in the written method, and they are the evidence that a small deviation in routine use does not require an investigation.

Downloads

Download the three-factor example workbook (.xlsx) — a three-way analysis with multiple crossed terms, main effect plots, and two- and three-way interaction plots: the analysis a robustness screen produces. A two-factor example shows the same with a single crossed term.

Common mistakes

Varying one factor at a time. It cannot detect interactions, and interactions are the failures that surprise people. Vary the factors together.

Choosing unrealistic levels. Extreme settings guarantee an effect and answer nothing. Use the deviation the method will actually meet.

Reading p-values instead of effect sizes. Statistical significance and practical importance are different questions. Judge the size of the effect against a pre-set limit.

Leaving interactions out of the model. Fitting main effects only pushes any interaction into the residual. That inflates the error term and hides the effect you were looking for.

Stopping at the finding. A sensitive parameter should end up with a tighter range in the written procedure, not a note in the validation report.

Run a robustness study with Analyse-it

Analyse-it fits the factorial model on your own runs, inside Excel:

  • Multi-factor models with crossed and interaction terms
  • The effect of each term, with main effect plots and two- and three-way interaction plots
  • Model fit diagnostics, so an effect is not read off a model that does not fit

Every feature from all five editions for 15 days, with no sign-up and no licence key. Model fitting is in every edition, from US$ 155 a year. Validated against NIST and CLSI reference datasets. See precision components explained for separating sources of variation in a nested rather than crossed design.