An effect size is a quantitative measure of the magnitude of a difference, association, or other phenomenon studied in statistics. The term can describe a population quantity or an estimate calculated from sample data. Effect sizes include differences between means, standardized differences, correlation coefficients, and ratios comparing event probabilities. Unlike a p-value, an effect size describes how large an observed relationship is rather than how incompatible the data are with a specified statistical model. Its meaning depends on the measure, outcome, and research context. (www2.psych.ubc.ca)
Magnitude and statistical significance
Effect size and statistical hypothesis testing answer different questions. A test evaluates evidence against a null hypothesis, whereas an effect-size estimate describes the magnitude and, where applicable, direction of a relationship. Statistical significance does not establish that an effect is substantively important: a small effect can produce a small p-value when estimated precisely. Conversely, an important effect may remain statistically inconclusive in an imprecise study. (amstat.org)
A population effect and its sample estimator must also be distinguished. An estimate varies across samples and can have bias. Reporting a confidence interval alongside it communicates uncertainty; a point estimate alone does not show how precisely the underlying effect has been determined. Nor does the word “effect” itself imply causation: a descriptive association requires additional design and identification assumptions before receiving a causal interpretation. (www2.psych.ubc.ca)
Differences between means
An unstandardized mean difference retains the outcome’s original units:
where each is a sample mean. A hypothetical difference of five examination points therefore remains interpretable on that examination’s scale. Such differences are especially useful when studies use the same measurement instrument. (cochrane.org)
A standardized mean difference expresses the difference relative to variability. For two independent groups, a common form of Cohen’s is
Here are sample standard deviations, and pools their variances. A hypothetical mean difference of five points with gives : the means differ by half a pooled standard deviation. The sign depends on the ordering of the groups. (frontiersin.org)
Hedges’ applies a correction to the upward bias of the magnitude of , particularly relevant in small samples. For paired or repeated measurements, different standardizing denominators produce different versions of standardized effects. Dividing by the standard deviation of change scores is not equivalent to dividing by the standard deviation of individual measurements; the chosen version therefore matters for comparisons across designs. (frontiersin.org)
Associations and explained variation
The Pearson coefficient measures linear association between two variables. It ranges from to ; its sign indicates direction, and its absolute magnitude indicates strength. A value near zero does not rule out a strong nonlinear relationship. (online.stat.psu.edu)
In simple linear regression fitted by ordinary least squares with an intercept, the coefficient of determination satisfies . It represents the proportion of observed outcome variation accounted for by the fitted model, not the proportion caused by the predictor. Regression slopes also describe effect magnitude, but their units depend on the scales of the predictor and outcome. (online.stat.psu.edu)
In analysis of variance, eta squared is
Partial eta squared instead divides the effect sum of squares by that effect’s sum of squares plus its associated error sum of squares. These quantities answer different questions and are not generally interchangeable, especially in multifactor or repeated-measures designs. (frontiersin.org)
Binary outcomes
For binary outcomes, let and denote event probabilities in two groups. Common measures are
The risk difference is an absolute comparison; the risk ratio and odds ratio are relative comparisons. The no-difference value is zero for , but one for and . (cochrane.org)
For hypothetical event probabilities of 0.20 and 0.10, these formulas give a difference of 10 percentage points, a risk ratio of 2, and an odds ratio of 2.25. An odds ratio is therefore not generally a risk ratio. Absolute and relative measures can convey different impressions because an absolute difference depends on the baseline event probability. (cochrane.org)
Interpretation and research synthesis
Conventional labels of “small,” “medium,” and “large”—often associated with values of 0.2, 0.5, and 0.8—are rough benchmarks, not universal boundaries of importance. Interpretation depends on outcome meaning, relevant comparisons, and the consequences of the observed difference. Standardization removes measurement units but does not remove contextual differences between populations. (frontiersin.org)
Effect sizes enter statistical power calculations alongside sample size, significance level, variability, and experimental design. Planning calculations may use an anticipated effect or a smallest effect of interest. A noisy pilot estimate is not necessarily a reliable planning value. (doi.org)
In meta-analysis, compatible effect estimates are combined, commonly using weights related to their precision. Standardized differences can support synthesis across instruments measuring the same construct, but differing population variability can alter their comparability. Reporting the measure, group ordering, denominator, standard error or interval, and design-specific calculation makes an estimate interpretable and reusable. A pooled effect also requires attention to study differences and biases rather than interpretation as a context-free constant. (cochrane.org)