One-hot encoding is a representation of a categorical value by a vector containing exactly one component equal to 1 and all other components equal to 0. Each position corresponds to a distinct category. In machine learning, it converts discrete labels into numerical inputs without assigning them an artificial order. In digital electronics, the same principle represents states using separate active bits. The number of components equals the number of categories or states represented. (developers.google.com)
Definition and mathematical structure
For a category set , the encoding maps to the vector , defined by
Each coordinate is an indicator function for membership in one category. For example, a vocabulary ordered as red, green, and blue gives:
| Category | Red | Green | Blue |
|---|---|---|---|
| Red | 1 | 0 | 0 |
| Green | 0 | 1 | 0 |
| Blue | 0 | 0 | 1 |
The category order fixes the meaning of each coordinate; it does not establish a ranking among categories. (developers.google.com)
Viewed within a real vector space, these vectors form the standard orthonormal basis of . Consequently, distinct categories have inner product zero and Euclidean distance . These mathematical consequences explain why the representation distinguishes identities but does not itself express semantic similarity: two closely related categories are initially as far apart as two unrelated ones. (developers.google.com)
Categorical features and statistical models
One-hot encoding is a common feature-engineering transformation for nominal variables, such as product type or geographic region. Assigning these categories consecutive integers can introduce unintended numerical relationships when a model treats the codes as quantities. One-hot encoding instead supplies a separate feature for each category, allowing category-specific model parameters. It is widely used with linear models and support vector machines. (developers.google.com)
For multiple categorical variables, their encoded vectors are concatenated into separate blocks. If the variables have categories, the resulting row has components and ordinarily ones. Thus, “one-hot” applies to each categorical block, not necessarily to the entire feature vector. Encoding a dataset produces a matrix whose columns retain fixed category meanings. (scikit-learn.org)
In linear regression with an intercept, retaining every indicator column creates a dependency: the columns for one categorical variable sum to the intercept column. The resulting design is not of full rank, so its coefficients are not uniquely identifiable without additional constraints. A common alternative retains indicators and treats the omitted category as a reference. This reduced representation is often called dummy coding; its reference category has an all-zero vector and is therefore not strictly one-hot. Dropping a category is not universally necessary, and it can affect models using regularization by breaking symmetry among categories. (scikit-learn.org)
Classification targets and embeddings
In supervised learning, one-hot vectors can represent mutually exclusive target classes. If and a model predicts class probabilities , the cross-entropy loss becomes
A softmax output commonly supplies these probabilities. Explicit target vectors are not always required: some loss implementations accept class indices and calculate the equivalent objective directly. Soft targets, including smoothed labels, can contain several nonzero probabilities and are not one-hot vectors. (docs.pytorch.org)
In natural language processing, vocabulary items can likewise be represented by one-hot vectors. A learned embedding transforms this sparse identity representation into a dense vector. If an embedding matrix stores one vector per row, then selects row . This algebraic interpretation underlies word embeddings; implementations can perform an indexed lookup without constructing the full one-hot input. The embedding’s learned geometry, unlike the original encoding, can capture relationships relevant to the training task. (developers.google.com)
Storage and vocabulary management
Large category vocabularies create high-dimensional representations. A dense dataset with observations and categories stores entries, although only entries are nonzero for one categorical variable. A sparse matrix stores nonzero values and their positions more efficiently. Sparse inputs, however, do not automatically eliminate the cost of a large model: connecting inputs to neural units still introduces weights. Embeddings and feature hashing offer alternative representations when vocabulary size is large. (developers.google.com)
An encoder requires a stable vocabulary and column order. Categories may be specified externally or learned from training data. During evaluation, learned preprocessing is fitted within the training portion and reused on held-out observations; fitting it using the test set risks data leakage. The same separation applies within cross-validation folds. (scikit-learn.org)
Previously unseen categories require an explicit policy. Implementations may reject them, map them to an unknown-category column, or return an all-zero block. Rare categories can be grouped into a shared bucket. An all-zero block is an operational extension rather than a strict one-hot vector; if a reference category was also dropped, the two cases may become indistinguishable. (scikit-learn.org)
Digital state encoding
In a finite-state machine, one-hot encoding assigns one storage bit to each state, with exactly one asserted during valid operation. A -state machine therefore uses state bits rather than the bits sufficient for compact binary encoding. This can simplify decoding and improve timing, at the cost of additional storage. Actual area and performance depend on the device and synthesized logic. All-zero and multiple-active-bit patterns are invalid under the strict scheme; recovery logic can explicitly return such patterns to a reset state. (intel.com)