Robustness testing A method that works when its developer runs it is not the same as a method that works. Robustness asks what happens when the incubation runs long, the pH drifts and a new reagent lot arrives — all at once.

Robustness testing varies the method’s parameters on purpose and checks that the result does not move more than you can tolerate. The changes are small, representing the drift that will happen anyway in routine use. Temperature, pH, flow rate, wavelength, incubation time, reagent lot, column batch: whichever of these the method is plausibly sensitive to. Analyst and day effects belong to intermediate precision instead; see the precision components.

ICH Q2(R2) and Q14 expect robustness to be evaluated, usually during development and before the validation study. Robustness is also where system suitability criteria come from. A parameter the method turns out to be sensitive to is one you have to control tightly and say so in the procedure. A parameter it tolerates is one you can stop worrying about.

One factor at a time answers the wrong question

The intuitive approach is to hold everything fixed, move one parameter, see what happens, put it back, and move the next. One factor at a time is inefficient, and worse, it cannot detect an interaction.

An interaction is where the effect of one factor depends on the level of another. Suppose the method tolerates a slightly high temperature and a slightly low pH, but recovery collapses at high temperature and low pH together. Varied one at a time, each change looks harmless, and the combination that breaks the method is never tested. Routine use varies several parameters at once, so that combination will eventually occur.

Two interaction plots of recovery against incubation temperature at two pH levels. On the left the two lines stay parallel, so each factor can be judged on its own. On the right they diverge: recovery holds at pH 7.4 but collapses at pH 7.0 when the temperature rises, so the effect of temperature depends on pH.
The signature of an interaction is non-parallel lines. Vary temperature alone and it looks harmless. Vary pH alone and it looks harmless. Only varying them together exposes the combination that breaks the method.

Vary them together

A factorial design tests every combination of the factor levels. With three factors at two levels each (low and high, either side of nominal), that is eight runs. Those eight runs estimate all three main effects and every interaction between them. To test the effects you also need an estimate of error: replicate the design, add centre points or treat the three-way interaction as error. Compared with one-factor-at-a-time, the factorial design gives more information from a similar number of runs. Every run contributes to the estimate of every effect rather than to just one.

Where there are many candidate parameters, a screening design tests a fraction of the combinations to identify which few factors matter. The cost is aliasing. In the smallest designs, such as Plackett–Burman, a main effect cannot be separated from two-factor interactions. Use at least a resolution IV design if two-factor interactions are a concern, and follow up any factor that shows an effect. The trade is usually right, because robustness is a screening question. The goal is to find the parameters that need controlling, not to characterise the response surface in detail.

Choose the levels to represent realistic deviation, not extremes. If the procedure says 37 °C and the incubator holds ±1 °C, test 36 and 38. Testing 30 and 45 will almost certainly find an effect, and will tell you nothing about whether the method is robust in service.

The analysis is a multi-factor model

Fit a model with each factor as a term and the crossed terms for their interactions. The model gives the effect of each factor on the result, the effect of each combination and a test of each against the residual variation. A main effects plot shows the size and direction of each factor’s influence. An interaction plot shows where the effect of one factor depends on the other. Non-parallel lines are the signature.

Read the effect sizes before the p-values. With enough replication a trivial effect becomes statistically significant. A significant effect that shifts the result by a tenth of your allowable variation is not a robustness problem. The question is not whether a factor has any effect, since everything has some effect. The question is whether that effect, over the realistic range, is large enough to matter. Judge it against a limit you set beforehand, from an allowable total error figure or an equivalent specification.

Conversely, a large effect that misses significance because the study was small is not evidence of robustness. Report the effect with its confidence interval and judge that interval against your limit, rather than treating a p-value as the answer.

An effect of terms table listing three factors A, B and C with every two-way and three-way interaction, each with its sum of squares, F statistic and p-value, followed by a grid of two-way interaction plots showing the response against each factor at both levels of another.
A three-factor design analysed in one model (surface finish, Montgomery). Every main effect and every interaction is estimated from the same sixteen runs. Factor A dominates at F = 18.69. B follows at 4.33 and the A×B interaction at 3.10, neither reaching significance in a study this size. The effect sizes therefore matter more than the p-values. The interaction plots below show the same information as shape. Parallel lines mean the effects simply add; converging or crossing lines mean they do not.

What to do with a factor that matters

A factor with an effect too large to tolerate does not mean the method has failed. Such a factor marks a parameter that needs a tighter specification, and the robustness study has just told you how tight. Narrow the permitted range until the effect across it falls inside your limit. Then write that range into the procedure and, where it can be monitored, into the system suitability criteria.

The factors that turn out not to matter are equally useful. Their results justify the tolerances in the written method, and show that a small deviation in routine use needs no investigation.

Downloads

Download the three-factor example workbook (.xlsx) — Montgomery’s surface-finish 2³ factorial, ready to open in the Analyse-it trial. Every two- and three-way interaction is fitted, with main effect plots and two- and three-way interaction plots. A robustness screen produces the same analysis. A two-factor example shows the analysis with a single crossed term.

Common mistakes

Varying one factor at a time. It cannot detect interactions, and interactions are the failures that surprise people. Vary the factors together.

Choosing unrealistic levels. Extreme settings guarantee an effect and answer nothing. Use the deviation the method will actually meet.

Reading p-values instead of effect sizes. Statistical significance and practical importance are different questions. Judge the size of the effect against a pre-set limit.

Leaving interactions out of the model. Fitting main effects only pushes any interaction into the residual, which inflates the error term and hides the effect you were looking for.

Stopping at the finding. A sensitive parameter should end up with a tighter range in the written procedure, not just a note in the validation report.

Run a robustness study with Analyse-it

Analyse-it fits the factorial model on your own runs, inside Excel:

  • Multi-factor models with crossed and interaction terms
  • The effect of each term, with main effect plots and two- and three-way interaction plots
  • Model fit diagnostics, so an effect is not read off a model that does not fit

Every feature from all five editions for 15 days. Model fitting is in every edition, from US$ 155 a year. Validated against NIST and CLSI reference datasets. See precision components explained for separating sources of variation in a nested rather than crossed design.