Process capability for non-normal data The capability formula assumes a bell curve. Feed it skewed data — flatness, roundness, particle counts, anything bounded at zero — and it will report a defect rate that can be out by an order of magnitude, in either direction.

Capability indices translate a process spread into an expected fraction outside the specification. That translation runs through the normal distribution. The calculation assumes the tails of the process fall away at the rate a bell curve predicts. When the real distribution is skewed or bounded, its tails behave differently, and an index built on them no longer means what it appears to mean.

Why skew breaks the usual index

Many real characteristics are not symmetric. Anything physically bounded at zero (a flatness, a concentration of an impurity, a time-to-event, a particle count) piles up against the bound and trails off to one side. Fit a symmetric normal curve to that shape and one tail is forced to describe data that are not there. The other understates the data that are.

Cpk is meant to estimate the tail beyond a limit. A mismatched tail is exactly the error that matters. A skewed process can look comfortably capable while still producing defects on its long side. Or look marginal when it is fine.

A right-skewed distribution against an upper specification limit: a fitted normal curve places its tail in the wrong position and misstates the fraction beyond the limit, while a transformed model matches the true skewed tail.
A symmetric normal curve forced onto skewed data puts its tail in the wrong place, misstating the fraction beyond the limit. Modelling the actual shape — typically by transforming the data to a scale where it is normal — corrects it.

Check normality before trusting the index

The first step is to look. A histogram with a normal overlay and a normal Q-Q plot will usually show skew at a glance. A formal normality test (Shapiro-Wilk or Anderson-Darling) puts a number on it. A clear departure from normality is the signal to stop reading the ordinary capability index at face value and to model the distribution properly. Do not skip this because the index “looks fine”. A healthy Cpk computed on skewed data is precisely the trap.

Transform, then compute the index

The practical route is to transform the data onto a scale where they are approximately normal. A Box-Cox or other power transformation finds the exponent that best straightens the distribution. Then compute capability on that scale, carrying the specification limits through the same transformation so that the index and the limits stay together.

The field recognises a second approach as well: fitting a non-normal distribution to the shape of the data and reading the out-of-specification fraction from it directly. Either route gives a defect estimate from a model that matches the tail you actually care about, rather than one imposed on it.

Report what you did

A capability index on transformed data is only interpretable with the transformation stated. Record the transformation used (the exponent, say) alongside the index. A reviewer needs to know that the 1.4 you are quoting came from a log scale, not the raw one. And that the specification limits were transformed to match, not left behind on the original scale.

Downloads

Download the process capability example (.xlsx) — includes a histogram, normal Q-Q plot, and Shapiro-Wilk normality test to judge the distribution before reading the index, ready to open in the Analyse-it trial.

Common mistakes

Assuming normality without checking. The index computes happily on any data; whether it means anything depends on the distribution. Test first.

Transforming the data but not the limits. Specification limits must go through the same transformation as the data, or the index compares two different scales.

Forcing normality on a bounded characteristic. A hard bound at zero guarantees skew. Model it rather than pretend it away.

Quoting a transformed index without saying so. An index means different things on different scales. State the transformation used.

Handle non-normal capability with Analyse-it

Analyse-it checks the distribution before it computes anything, inside Excel:

  • A histogram, a normal Q-Q plot with a Lilliefors band, and Shapiro-Wilk, Anderson-Darling and Kolmogorov-Smirnov tests
  • Box-Cox and power transformations where the data are not normal
  • Capability computed after that, rather than on an assumption nobody checked

Every feature from all five editions for 15 days, with no sign-up and no licence key. Capability analysis is in the Quality Control & Improvement and Ultimate editions, from US$ 290 a year. Validated against published reference datasets and thousands of internal test cases. See Cp, Cpk, Pp and Ppk explained for the indices themselves.