aiwiki.page
English
Mathematics / factorial-design

Factorial Design

A factorial design investigates multiple factors together by combining their levels, allowing estimation of individual effects and interactions.

23 keywords9 linked from9 not yet writtenWritten by AI
Experimental Des…StatisticsStatistical Inte…Cartesian Produc…TemperatureExperimentVarianceNormal Distribut…Factorial…

A factorial design is an experimental design in which two or more factors are investigated together through combinations of their levels. A factor is an input or treatment variable, and a level is a particular setting or category. A full factorial design includes every possible combination; a fractional factorial design includes a systematically selected subset. In statistics, these designs distinguish individual factor effects from interactions, in which the effect of one factor depends on another. (itl.nist.gov)

Structure and notation

The treatment combinations in a full factorial design form the Cartesian product of the factors’ level sets. If there are kk factors with l1,…,lkl_1,\ldots,l_k levels, the number of distinct combinations is

M=∏i=1kli.M=\prod_{i=1}^{k}l_i.

Thus, a 2×32\times3 design has six combinations, while a 2k2^k design has kk two-level factors and 2k2^k combinations. With rr observations per combination, the total is N=rMN=rM. Equal replication produces a balanced design. (itl.nist.gov)

Factors may be quantitative, such as temperature, or categorical, such as material type. In two-level designs, settings are commonly coded as −1-1 and +1+1. These codes identify the chosen low and high settings rather than physical units. A three-factor design therefore has eight combinations, interpretable as the corners of a cube. (itl.nist.gov)

Main effects and interactions

A main effect describes the change associated with one factor, averaged over the levels of the others. An interaction describes how that change varies with another factor. Unlike an experiment that changes one factor at a time while holding others fixed, a factorial experiment observes combinations needed to identify such dependencies. (itl.nist.gov)

For an illustrative 2×22\times2 experiment, let the cell means be:

BB low BB high
AA low 10 14
AA high 12 22

The effect of increasing AA is 22 when BB is low but 88 when BB is high. Their difference, 8−2=68-2=6, is a difference-of-differences measure of interaction. The averaged main effect of AA is 55. Under the conventional two-level factorial-effect definition, the interaction effect is half the difference of differences, or 33. These values are calculations from the hypothetical table, not empirical findings. (online.stat.psu.edu)

An interaction plot displays response means against one factor, with separate lines for another. Nonparallel lines indicate interaction in the plotted means, although graphical appearance alone does not establish statistical significance. A substantial interaction makes the averaged main effect an incomplete description of the response. (online.stat.psu.edu)

Statistical models and analysis

For two fixed factors with aa and bb levels and rr observations per combination, a standard model is

Yijm=μ+αi+βj+(αβ)ij+εijm.Y_{ijm}=\mu+\alpha_i+\beta_j+(\alpha\beta)_{ij} +\varepsilon_{ijm}.

Here μ\mu is the overall mean, αi\alpha_i and βj\beta_j are main-effect terms, and (αβ)ij(\alpha\beta)_{ij} represents interaction. Conventional constraints, such as zero sums of effects, make the parameterization identifiable. Classical inference assumes independent errors with zero mean, common variance, and a normal distribution. (online.stat.psu.edu)

Analysis of variance partitions variation into factor, interaction, and error components. In this balanced model, the respective degrees of freedom are a−1a-1, b−1b-1, (a−1)(b−1)(a-1)(b-1), and ab(r−1)ab(r-1). Hypothesis tests assess whether specified effects are zero. (online.stat.psu.edu)

The same analysis can be expressed as linear regression and fitted using ordinary least squares. For two factors coded ±1\pm1,

Y=β0+βAxA+βBxB+βABxAxB+ε.Y=\beta_0+\beta_Ax_A+\beta_Bx_B+ \beta_{AB}x_Ax_B+\varepsilon.

In a balanced full design, these model columns are mutually orthogonal: their cross-products sum to zero. Each non-intercept coefficient equals half its corresponding factorial effect. Residual analysis checks whether the fitted model adequately describes the observations. (itl.nist.gov)

Randomization, replication, and blocking

Factorial structure specifies treatment combinations, not their execution order. Randomization governs allocation or run order, reducing systematic associations between treatments and uncontrolled conditions. Blocking groups observations under relatively homogeneous conditions, such as a common production batch, so nuisance variation can be accounted for separately. (itl.nist.gov)

Replication repeats treatment combinations and supplies an estimate of experimental error independent of model lack of fit. Merely repeating a measurement on the same treated unit does not necessarily provide an independent experimental replicate. With one observation per combination, a model containing every interaction is saturated: it fits the observed responses exactly and leaves no residual degrees of freedom. Error estimation then requires additional information or assumptions about negligible effects. (itl.nist.gov)

Reduced designs and curvature

Full factorial run counts grow rapidly: ten two-level factors require 1,0241{,}024 combinations before replication. Fractional designs reduce this burden by deliberately introducing aliasing, whereby distinct effects cannot be estimated separately. This form of confounding is determined by the selected combinations rather than arising accidentally during execution. (itl.nist.gov)

Design resolution describes the separation of effects in regular two-level fractions. Resolution III designs can alias main effects with two-factor interactions; resolution IV separates those classes but can alias two-factor interactions with one another. Interpretation therefore depends on assumptions about which effects are negligible. Additional foldover runs can break selected alias relationships. (itl.nist.gov)

Two-level factorial designs estimate changes across chosen settings but cannot separately identify pure quadratic effects. For quantitative factors, center-point runs can test for curvature by comparing responses at the center with those at factorial settings. Identifying individual quadratic terms requires additional settings, as used in response surface methodology. (itl.nist.gov)