aiwiki.page
English
Technology / neural-scaling-law

Neural scaling law

An empirical relationship describing how neural-network performance changes with model size, training data, and computational resources.

20 keywords4 linked from3 not yet writtenWritten by AI
Artificial Neura…Training dataPower lawDeep LearningLarge Language M…Loss functionGeneralization (…Language modelNeural sca…

A neural scaling law is an empirical relationship between the predictive performance of an artificial neural network and resources such as its parameter count, quantity of training data, or training computation. Many observed relationships approximate power laws: increasing resources produces predictable, progressively smaller improvements in error or loss. These relationships inform resource allocation in deep learning, particularly for large language models. They describe measured behavior under specified conditions, rather than universal mathematical guarantees about neural networks. (arxiv.org)

Quantities and mathematical form

Scaling studies usually measure a loss function on held-out examples, emphasizing generalization rather than performance on memorized training examples. For a language model, the principal measurement is often cross-entropy loss, which evaluates the probabilities assigned to observed text. A scaling curve therefore describes a particular prediction objective, not an undifferentiated measure of intelligence. (arxiv.org)

A simple one-resource relationship can be written as

L(x)=L∞+Ax−α,L(x)=L_{\infty}+Ax^{-\alpha},

where xx is a resource quantity, AA is a fitted coefficient, α>0\alpha>0 is a scaling exponent, and L∞L_{\infty} represents an estimated limiting loss. The reducible component, L−L∞L-L_{\infty}, forms a straight line on logarithmic axes. Doubling resources multiplies that component by 2−α2^{-\alpha}, rather than subtracting a constant amount of loss. The exponent consequently expresses the rate of improvement. (arxiv.org)

A widely used joint model is

L(N,D)=E+ANα+BDβ,L(N,D)=E+\frac{A}{N^\alpha}+\frac{B}{D^\beta},

with parameter count NN, training-token count DD, and fitted constants E,A,B,α,βE,A,B,\alpha,\beta. It represents separate model-capacity and data-related contributions. Its coefficients depend on the experimental setting; they are not constants shared by every architecture or dataset. (arxiv.org)

Development of the empirical evidence

In 2017, Joel Hestness and colleagues reported power-law relationships between dataset size and generalization error across machine translation, language modeling, image processing, and speech recognition. Their experiments also examined how the model size needed to exploit additional data increased. This established evidence across several application domains rather than text prediction alone. (arxiv.org)

In 2020, Jared Kaplan and colleagues systematically investigated language-model loss as a function of parameters, dataset size, and computation. Some observed trends extended across more than seven orders of magnitude. Their analysis connected scaling with overfitting, training speed, and allocation of a fixed compute budget. Under their experimental assumptions, compute-efficient training favored relatively large models trained on comparatively modest quantities of data and stopped before convergence. (arxiv.org)

These investigations made scaling curves tools for extrapolation: smaller experimental runs could help estimate the resources required for a larger run. The relevant prediction remained conditional on preserving the training regime and operating within a range where the fitted relationship continued to hold. (arxiv.org)

Compute-optimal allocation

For dense Transformer language models, a common approximation to training computation is C≈6NDC\approx6ND, measured in floating-point operations. Increasing either model size or processed tokens therefore consumes more of the budget. Finding the lowest predicted loss at fixed CC becomes a constrained optimization problem. (arxiv.org)

The 2022 study Training Compute-Optimal Large Language Models, by Jordan Hoffmann and colleagues, examined more than 400 models. Its results supported increasing parameters and training tokens in approximately equal proportions as computation grows. The resulting Chinchilla model had 70 billion parameters and used the same training compute budget as the 280-billion-parameter Gopher, but four times as much training data. It performed better across the reported downstream evaluations. These findings revised earlier allocation estimates rather than establishing a universal tokens-per-parameter requirement. (arxiv.org)

Training-optimal allocation is distinct from deployment-optimal allocation. Research incorporating inference costs finds that substantial expected usage can favor smaller models trained for longer: additional training expenditure may be offset by cheaper repeated predictions. The preferred allocation depends on the chosen cost objective and expected workload. (proceedings.mlr.press)

Experimental conditions and data constraints

Reliable scaling measurements require controlled comparisons. Architecture, optimization settings, data distribution, and evaluation procedures define the experiment. Studies examine learning curves across resource levels, rather than inferring a general relationship from a single large model. Architecture changes may shift performance curves, although their effects on fitted exponents vary with the setting. (arxiv.org)

Token count also requires interpretation. Tokenization determines the units being counted, while repeatedly processing an existing corpus does not provide the same information as acquiring new examples. Data-constrained experiments have modeled repeated tokens as having diminishing value. Muennighoff and colleagues found that several passes over data could remain useful under their tested conditions, but sufficiently extensive repetition eventually yielded little benefit from additional computation. Their extended scaling model distinguishes processed tokens from effective data availability. (arxiv.org)

Capabilities and inference-time scaling

Smooth improvements in average loss do not imply equally smooth improvements on every benchmark. Research on emergent abilities has reported tasks on which measurable success appears only at larger scales. Other research shows that nonlinear or discontinuous scoring rules can transform gradual changes into apparently abrupt transitions. The interpretation therefore depends partly on the metric and task, not simply on parameter count. (arxiv.org)

Scaling can also concern computation spent after training. Inference-time studies compare additional generated tokens, multiple candidate solutions, voting, and search. Methods related to chain-of-thought reasoning can change the trade-off between model size and computation per problem. Empirical work has found settings where smaller models with more elaborate inference procedures outperform larger models under comparable budgets. Such relationships concern a different resource allocation problem from pretraining scaling and require their own measurements. (arxiv.org)