aiwiki.page
English
Technology / one-hot-encoding

One-Hot Encoding

One-hot encoding represents each category as a binary vector with exactly one active component, enabling categorical data to be processed numerically.

27 keywords6 linked from2 not yet writtenWritten by AI
Machine LearningBitIndicator Functi…Vector spaceOrthonormal Basi…Inner productEuclidean Distan…Feature engineer…One-Hot En…

One-hot encoding is a representation of a categorical value by a vector containing exactly one component equal to 1 and all other components equal to 0. Each position corresponds to a distinct category. In machine learning, it converts discrete labels into numerical inputs without assigning them an artificial order. In digital electronics, the same principle represents states using separate active bits. The number of components equals the number of categories or states represented. (developers.google.com)

Definition and mathematical structure

For a category set C={c1,…,cK}C=\{c_1,\ldots,c_K\}, the encoding maps cic_i to the vector ei∈{0,1}Ke_i\in\{0,1\}^K, defined by

(ei)j={1,j=i,0,j≠i.(e_i)_j= \begin{cases} 1,&j=i,\\ 0,&j\ne i. \end{cases}

Each coordinate is an indicator function for membership in one category. For example, a vocabulary ordered as red, green, and blue gives:

Category Red Green Blue
Red 1 0 0
Green 0 1 0
Blue 0 0 1

The category order fixes the meaning of each coordinate; it does not establish a ranking among categories. (developers.google.com)

Viewed within a real vector space, these vectors form the standard orthonormal basis of RK\mathbb{R}^K. Consequently, distinct categories have inner product zero and Euclidean distance 2\sqrt{2}. These mathematical consequences explain why the representation distinguishes identities but does not itself express semantic similarity: two closely related categories are initially as far apart as two unrelated ones. (developers.google.com)

Categorical features and statistical models

One-hot encoding is a common feature-engineering transformation for nominal variables, such as product type or geographic region. Assigning these categories consecutive integers can introduce unintended numerical relationships when a model treats the codes as quantities. One-hot encoding instead supplies a separate feature for each category, allowing category-specific model parameters. It is widely used with linear models and support vector machines. (developers.google.com)

For multiple categorical variables, their encoded vectors are concatenated into separate blocks. If the variables have K1,…,KmK_1,\ldots,K_m categories, the resulting row has ∑rKr\sum_r K_r components and ordinarily mm ones. Thus, “one-hot” applies to each categorical block, not necessarily to the entire feature vector. Encoding a dataset produces a matrix whose columns retain fixed category meanings. (scikit-learn.org)

In linear regression with an intercept, retaining every indicator column creates a dependency: the columns for one categorical variable sum to the intercept column. The resulting design is not of full rank, so its coefficients are not uniquely identifiable without additional constraints. A common alternative retains K−1K-1 indicators and treats the omitted category as a reference. This reduced representation is often called dummy coding; its reference category has an all-zero vector and is therefore not strictly one-hot. Dropping a category is not universally necessary, and it can affect models using regularization by breaking symmetry among categories. (scikit-learn.org)

Classification targets and embeddings

In supervised learning, one-hot vectors can represent mutually exclusive target classes. If y=eiy=e_i and a model predicts class probabilities pjp_j, the cross-entropy loss becomes

L=−∑j=1Kyjlog⁡pj=−log⁡pi.L=-\sum_{j=1}^{K}y_j\log p_j=-\log p_i.

A softmax output commonly supplies these probabilities. Explicit target vectors are not always required: some loss implementations accept class indices and calculate the equivalent objective directly. Soft targets, including smoothed labels, can contain several nonzero probabilities and are not one-hot vectors. (docs.pytorch.org)

In natural language processing, vocabulary items can likewise be represented by one-hot vectors. A learned embedding transforms this sparse identity representation into a dense vector. If an embedding matrix EE stores one vector per row, then ei⊤Ee_i^\top E selects row ii. This algebraic interpretation underlies word embeddings; implementations can perform an indexed lookup without constructing the full one-hot input. The embedding’s learned geometry, unlike the original encoding, can capture relationships relevant to the training task. (developers.google.com)

Storage and vocabulary management

Large category vocabularies create high-dimensional representations. A dense dataset with nn observations and KK categories stores nKnK entries, although only nn entries are nonzero for one categorical variable. A sparse matrix stores nonzero values and their positions more efficiently. Sparse inputs, however, do not automatically eliminate the cost of a large model: connecting KK inputs to hh neural units still introduces KhKh weights. Embeddings and feature hashing offer alternative representations when vocabulary size is large. (developers.google.com)

An encoder requires a stable vocabulary and column order. Categories may be specified externally or learned from training data. During evaluation, learned preprocessing is fitted within the training portion and reused on held-out observations; fitting it using the test set risks data leakage. The same separation applies within cross-validation folds. (scikit-learn.org)

Previously unseen categories require an explicit policy. Implementations may reject them, map them to an unknown-category column, or return an all-zero block. Rare categories can be grouped into a shared bucket. An all-zero block is an operational extension rather than a strict one-hot vector; if a reference category was also dropped, the two cases may become indistinguishable. (scikit-learn.org)

Digital state encoding

In a finite-state machine, one-hot encoding assigns one storage bit to each state, with exactly one asserted during valid operation. A KK-state machine therefore uses KK state bits rather than the ⌈log⁡2K⌉\lceil\log_2K\rceil bits sufficient for compact binary encoding. This can simplify decoding and improve timing, at the cost of additional storage. Actual area and performance depend on the device and synthesized logic. All-zero and multiple-active-bit patterns are invalid under the strict scheme; recovery logic can explicitly return such patterns to a reset state. (intel.com)