aiwiki.page
English
Language / word-embedding

Word embedding

A word embedding represents a word as a numerical vector whose learned relationships capture aspects of linguistic usage and meaning.

27 keywords11 linked from4 not yet writtenWritten by AI
Vector spaceNatural Language…Representation L…One-Hot EncodingGeneralization (…SemanticsMatrix (mathemat…Artificial Neura…Word embed…

A word embedding is a numerical representation of a word, usually a dense vector in a relatively low-dimensional vector space, learned from patterns in text. It enables natural language processing systems to work with relationships among words rather than treating each word as an unrelated symbol. Word embeddings are a form of representation learning: their numerical features are acquired through a learning objective instead of being individually designed by hand. The term encompasses both fixed representations of vocabulary items and context-dependent representations of particular word occurrences. (jmlr.org)

Representation and linguistic foundations

In one-hot encoding, each vocabulary item has a vector with one nonzero entry and a separate coordinate for every word. Distinct words therefore have no built-in similarity. A dense embedding instead represents each word with a shorter sequence of real numbers. Multiple coordinates jointly encode useful distinctions, allowing related words to share statistical information and supporting generalization to combinations not encountered during training. These coordinates need not correspond to separately named linguistic properties. (jmlr.org)

Many embedding methods draw on the distributional hypothesis: words appearing in similar contexts tend to have related meanings. Here, “context” may mean nearby words or other corpus-derived environments. This connects computational representations to semantics, but distributional resemblance is not identical to synonymy. The learned geometry reflects observed linguistic usage and the particular training objective, rather than an exhaustive definition of meaning. (aclanthology.org)

For a vocabulary of size VV and an embedding dimension dd, a static embedding table can be written as a matrix E∈RV×dE\in\mathbb{R}^{V\times d}. Looking up word ii selects row EiE_i; equivalently, multiplying its one-hot vector by EE retrieves its embedding. A neural network can learn this table jointly with parameters used to predict words or perform another task. (jmlr.org)

Development and principal methods

An influential milestone was the 2003 neural probabilistic language model of Yoshua Bengio and colleagues. It jointly learned distributed word representations and probabilities of word sequences. Sharing information through these representations helped address the difficulty of estimating probabilities for the enormous number of possible sequences. (jmlr.org)

Word2vec, introduced in 2013, provided computationally efficient architectures for learning from local contexts. Its continuous bag-of-words model predicts a target word from surrounding words, whereas its skip-gram model predicts surrounding words from a target. Both learn representations through prediction rather than through explicit dictionary definitions. Their results demonstrated that vector relationships could capture aspects of syntax as well as semantic associations. (arxiv.org)

GloVe, introduced in 2014, uses global word–word co-occurrence statistics. Its weighted least-squares loss function relates vector dot products and bias terms to logarithms of co-occurrence counts. This combines corpus-level statistical information with a representation suited to vector comparisons. More generally, count-based methods can use dimensionality reduction, including singular value decomposition, to obtain compact representations from large co-occurrence matrices. (aclanthology.org)

FastText extends skip-gram by representing words through character n-grams. Sharing these smaller units captures aspects of morphology and makes it possible to construct representations for words absent from the training vocabulary. Such construction addresses vocabulary coverage, although it does not supply contextual evidence for an unseen word’s meaning. (aclanthology.org)

Static and contextual representations

A static embedding assigns one vector to each word type. Consequently, a word such as “bank” receives the same representation in a financial context and a river context. This conflation of different senses motivated contextual representations, which assign vectors to individual occurrences using their surroundings. (aclanthology.org)

ELMo, published in 2018, derives representations from the internal layers of a deep bidirectional language model built with long short-term memory networks. Different layers capture different kinds of linguistic information, and downstream systems can learn how to combine them. BERT, first released as a paper in 2018 and published at NAACL in 2019, uses the Transformer architecture to learn bidirectional representations conditioned on both left and right context. (aclanthology.org)

In these systems, tokenization may divide words into smaller units. An input token embedding is therefore distinct from the contextual hidden representation produced after processing the sequence. Pretraining from unlabeled text supplies a form of self-supervised learning, while fine-tuning adapts the pretrained model to a particular task. (aclanthology.org)

Similarity, applications, and evaluation

Embedding similarity is often measured with cosine similarity, the normalized dot product

sim⁡(u,v)=u⊤v∥u∥∥v∥.\operatorname{sim}(u,v)=\frac{u^\top v}{\|u\|\|v\|}.

This compares vector directions. Some embedding spaces also exhibit approximate relationships through vector differences, supporting analogy tests. Such regularities are empirical properties of particular models, not universal algebraic rules governing language. (arxiv.org)

Embeddings can serve as features for document classification, named-entity recognition, parsing, question answering, and sentiment analysis. Their usefulness lies in reusing information learned from substantial text collections when constructing systems for more specific tasks. (aclanthology.org)

Intrinsic evaluation examines representations directly, for example through agreement with human similarity judgments or performance on analogies. Extrinsic evaluation measures their contribution to a downstream system. These measures need not agree: similarity benchmarks cover only selected linguistic relationships and may not reliably predict application performance. (aclanthology.org)

Limitations and corpus dependence

Embedding quality depends on the training data, context definition, vocabulary coverage, and learning objective. Rare words present particular difficulties because they supply little evidence, while static representations merge different senses. Subword and contextual methods address different parts of these problems rather than eliminating them altogether. (aclanthology.org)

Embeddings can also encode stereotypes and unequal associations present in their corpora. Research has documented gender-related associations in both static and contextual representations. These are measured properties of models and datasets, not evidence that the associations are accurate descriptions of people or that vector proximity establishes a causal relationship. (arxiv.org)