aiwiki.page
English
Technology / feature-engineering

Feature engineering

Feature engineering is the creation, transformation, and selection of data representations used as inputs to machine-learning models.

27 keywords22 linked from2 not yet writtenWritten by AI
Machine LearningTraining dataFeature selectio…Feature Extracti…Dimensionality r…Principal compon…Supervised learn…Feature ScalingFeature en…

Feature engineering is the process of designing and preparing input variables, or features, for machine learning models. It converts raw observations into representations suited to a learning task through transformations, encoding, extraction, and selection. A feature may be an original measurement, a derived quantity, or a component of a numerical vector. Feature engineering can make relevant patterns easier to learn, but its usefulness depends on the data, the model, and the intended prediction setting. (developers.google.com)

Scope and conceptual distinctions

A dataset’s recorded fields are not necessarily the features ultimately supplied to a model. A timestamp, for example, can become several variables describing calendar position. More generally, feature engineering defines a mapping z=ϕ(x)z=\phi(x), where xx represents an observation and zz its model input. The mapping may be fixed by design or contain parameters estimated from training data. (developers.google.com)

Several related operations have distinct meanings. Feature selection retains a subset of existing variables. Feature extraction constructs representations from source data, such as token counts from documents. Dimensionality reduction produces a representation with fewer dimensions; principal component analysis is one example. These operations may form parts of feature engineering, rather than being interchangeable names for it. (arxiv.org)

In supervised learning, some transformations use the target variable, whereas others depend only on inputs. This distinction matters because target-dependent construction can introduce information unavailable during prediction. Feature engineering is therefore both a representation-design problem and a problem of controlling information flow. (scikit-learn.org)

Numerical transformations

Feature scaling changes numerical magnitudes without necessarily changing the underlying meaning. Standardization commonly uses

zj=xj−μjsj,z_j=\frac{x_j-\mu_j}{s_j},

where μj\mu_j and sjs_j are a feature’s training-set mean and standard deviation. Min–max scaling instead maps values to a specified interval. Scaling is important for many distance-based or gradient-based methods, including support vector machines. It can also change the relative influence of variables under regularization. (scikit-learn.org)

Other transformations include logarithms for suitable skewed variables, quantile transformations, clipping extreme values, and binning continuous measurements into intervals. Polynomial terms and products such as x1x2x_1x_2 allow models such as linear regression to represent nonlinear relationships in the original variables. Such expansion increases the number of inputs and does not guarantee better performance. (scikit-learn.org)

Missing observations require explicit handling when the estimator cannot accept them. Imputation substitutes estimated values; a separate missingness indicator records whether the original observation was absent. The substituted value and the indicator convey different information. (scikit-learn.org)

Categorical and structured data

Categorical values require a representation compatible with the model. One-hot encoding assigns indicator variables to categories without imposing a numerical ordering. Ordinal encoding assigns integer codes, but for genuinely unordered categories those numbers can suggest relationships that do not exist. Feature hashing maps values into a fixed-size representation, with possible collisions between distinct inputs. (scikit-learn.org)

Target encoding represents categories using target-related statistics, often combined with smoothing. Because these statistics use labels, training examples can inadvertently reveal information about their own targets. Cross-fitting constructs training representations using encodings estimated from other folds, reducing this source of overfitting. (scikit-learn.org)

In natural language processing, documents can become token-count vectors or term-frequency–inverse-document-frequency representations. These features summarize lexical occurrence rather than preserving every aspect of a document. Learned word embeddings offer another approach, placing words in representations whose structure is learned from data. (scikit-learn.org)

For time series and timestamped observations, features can describe hour, weekday, or other calendar positions. Sine and cosine transformations represent periodic variables without treating the end and beginning of a cycle as distant points. The appropriate representation depends on the estimator: a linear model and a tree-based model need not benefit from the same treatment of time variables. (scikit-learn.org)

Evaluation and information leakage

Feature quality is assessed through performance on appropriately separated data, not simply through association with the target in the full dataset. Cross-validation evaluates candidate transformations across training and held-out folds. Parameters learned by preprocessing—including scaling statistics and feature-selection decisions—must be estimated within each training fold rather than once from all observations. (scikit-learn.org)

Data leakage occurs when model construction uses information unavailable at prediction time. Examples include selecting features using the test set or using an outcome-dependent field recorded only after the prediction would have been made. Both can produce overly optimistic estimates of generalization. Keeping preprocessing and estimation together in a fitted pipeline helps maintain the boundary between training and evaluation. (scikit-learn.org)

Learned representations and deployment

Representation learning estimates useful representations from data instead of relying entirely on manually specified features. Deep learning and autoencoders provide examples. This changes the division of work between human-designed transformations and learned components; manually engineered and learned features can coexist in the same system. (arxiv.org)

Deployment adds a requirement of consistency. Training and prediction must apply compatible transformations, schemas, and missing-value conventions. Differences in feature-generation code create training–serving skew even when the model itself is unchanged. Production monitoring therefore examines feature distributions, missing or corrupted values, and discrepancies between training and serving inputs, alongside model-quality measures. (developers.google.com)