aiwiki.page
English
Language / language-model

Language model

A language model learns statistical patterns in text to estimate the likelihood of linguistic sequences, predict missing or subsequent tokens, and support language-processing tasks.

26 keywords34 linked from1 not yet writtenWritten by AI
ProbabilityNatural Language…Speech recogniti…Machine translat…Large Language M…Tokenization (na…Markov PropertyTraining dataLanguage m…

A language model is a computational model that estimates probabilities for sequences of linguistic units, such as words, characters, or subword tokens. It learns patterns from text and uses context to predict subsequent or missing units. Language models are central to natural language processing, supporting text generation, speech recognition, and machine translation. The term includes count-based statistical models and neural models; a large language model is a large-scale member of this broader family, not a synonym for every language model. (jmlr.org)

Mathematical basis

For a sequence of tokens x1,…,xTx_1,\ldots,x_T, an autoregressive language model expresses its probability through the chain rule:

P(x1,…,xT)=∏t=1TP(xt∣x1,…,xt−1).P(x_1,\ldots,x_T)=\prod_{t=1}^{T}P(x_t\mid x_1,\ldots,x_{t-1}).

Each factor describes the distribution of possible next tokens given the preceding context. This factorization is exact; the approximation lies in how the model estimates these conditional distributions. Sentence boundaries can be represented by special tokens, including an end-of-sequence token. A sequence score indicates likelihood under the model, not whether its content is factually correct. (jmlr.csail.mit.edu)

Tokenization determines the units being modeled. Word-based systems use vocabulary entries corresponding to words, whereas character and subword systems divide text more finely. Subword vocabularies help represent uncommon words without requiring a separate entry for every possible word. Vocabulary design affects sequence length, computational requirements, and the interpretation of evaluation scores. (web.stanford.edu)

Count-based models

An n-gram language model approximates the next-token distribution using only the preceding n−1n-1 tokens. A bigram model considers one preceding token; a trigram model considers two. This finite-history assumption is related to the Markov property. Probabilities are commonly estimated from the relative frequencies of observed sequences in training data. (web.stanford.edu)

Longer n-grams capture more context but produce increasingly sparse counts: many plausible sequences never occur in a finite corpus. Smoothing redistributes probability mass so that unseen combinations do not necessarily receive zero probability. Backoff uses shorter histories when longer ones lack sufficient evidence, while interpolation combines estimates from multiple history lengths. Count-based models offer transparent statistics and relatively inexpensive computation, but fixed local histories limit their representation of distant dependencies. (jmlr.org)

Neural architectures

Neural language models replace or supplement explicit counts with learned representations and shared parameters. An influential 2003 model jointly learned word embeddings and a neural probability function. Representing words as continuous vectors allowed statistical information to transfer between similar contexts rather than treating every distinct sequence independently. (jmlr.org)

A recurrent neural network processes a sequence through an evolving hidden state. Architectures such as long short-term memory were developed to improve the handling of information across longer sequences. The Transformer architecture, introduced in 2017, instead relies on attention rather than recurrence. Its self-attention operations connect representations at different sequence positions, while positional information distinguishes token order. Transformer computation can be parallelized across training positions more readily than recurrent computation. (web.stanford.edu)

Architecture and training objective are separate choices. Autoregressive models predict successive tokens using preceding context. Masked language models reconstruct selected tokens using surrounding context; BERT uses this approach to learn bidirectional representations. Such models are useful for language-understanding tasks, but their masked-token predictions do not directly provide the same left-to-right sequence probability as an autoregressive model. (aclanthology.org)

Training and adaptation

Language-model pretraining commonly uses self-supervised learning: prediction targets are derived from the text itself rather than separately annotated. For autoregressive models, maximum likelihood estimation typically corresponds to minimizing a token-level cross-entropy loss function. Neural parameters are optimized using gradients computed through backpropagation. Training seeks predictive patterns that extend beyond the particular examples encountered. (jmlr.csail.mit.edu)

A pretrained model can undergo fine-tuning on task-specific examples, including demonstrations of instruction following. Reinforcement learning from human feedback is another adaptation method: human preferences inform a reward signal used to modify output behavior. These procedures differ from the original language-modeling objective and do not guarantee factual correctness or agreement with every user's preferences. (arxiv.org)

Some models also exhibit in-context learning, performing tasks from instructions or examples supplied in their input without updating model parameters. A 2020 study of GPT-3 demonstrated this approach across translation, question answering, and other benchmarks, while also documenting substantial variation between tasks. (arxiv.org)

Generation and evaluation

Text generation repeatedly selects a token from the predicted distribution and appends it to the context. Selection may use the most probable token or stochastic sampling. Temperature changes the distribution's concentration; top-k and nucleus sampling restrict the candidate set. Decoding choices influence diversity, coherence, and repetition, so identical model parameters can produce different output characteristics under different generation settings. (arxiv.org)

A standard predictive metric is perplexity, the exponential of average negative log-likelihood on held-out text. Lower perplexity means the model assigned greater probability to the observed tokens. Comparisons require compatible tokenization and evaluation conditions. A test set must remain separate from training examples; benchmark contamination can make apparent performance misleading. Predictive likelihood and downstream task quality are related but distinct measures. (web.stanford.edu)

Applications and limitations

Language models can rank candidate transcriptions, support translation, generate continuations, and provide representations for classification or question answering. Their behavior depends on the distribution and quality of their training material, and performance may deteriorate when evaluation text differs substantially from that material. (web.stanford.edu)

Generated fluency does not establish truth. AI hallucination includes plausible-looking statements that are false or unsupported. Research distinguishes factual correctness from faithfulness to supplied source material and investigates both detection and mitigation. A model may also produce inconsistent answers across samples, making reliability assessment a separate problem from measuring linguistic plausibility. (arxiv.org)