Feature extraction is the process of transforming data into measurable attributes or representations that a computational system can use. In machine learning, these features provide inputs for prediction, classification, or other analytical tasks. An extractor may convert text into numerical vectors, describe an image through local patterns, or project existing measurements onto new coordinates. The transformation may be specified by human designers or learned from data; its purpose is to expose useful structure in a form compatible with subsequent processing. (scikit-learn.org)
Scope and related concepts
A feature extractor can be expressed as a mapping , where is an observation, its representation, and denotes any adjustable parameters. Some mappings have no learned parameters; others derive their parameters from examples. A downstream predictor then operates on , rather than directly on the original input. This separation can be explicit, as in a preprocessing pipeline, or internal to a jointly trained neural network. (deeplearningbook.org)
Feature extraction differs from feature selection: selection retains a subset of existing attributes, whereas extraction constructs or encodes a representation. Feature engineering is a broader practical activity encompassing the design, transformation, and organization of features. Dimensionality reduction overlaps with extraction when the resulting representation has fewer coordinates, but extraction does not necessarily reduce dimensionality. A text document, for example, may become a vector with thousands of vocabulary coordinates. Representation learning specifically concerns discovering representations from data rather than defining all their properties manually. (scikit-learn.org)
Designed representations
In natural language processing, the bag-of-words model represents documents through token occurrence counts, usually disregarding their order. Tokenization determines the units being counted; these may be words, subwords, or other units. Term frequency–inverse document frequency, or TF–IDF, adjusts counts using information about how widely terms occur across a document collection. Such representations are commonly stored in a sparse matrix because most documents contain only a small fraction of the vocabulary. Their simplicity makes them computationally useful, although ignoring order limits the linguistic information they retain. (scikit-learn.org)
Feature hashing maps symbolic feature names into a fixed number of coordinates through a hash function. It avoids maintaining a complete feature-name dictionary and can accommodate previously unseen names. However, different features may map to the same coordinate, producing collisions. The chosen representation size therefore affects both storage requirements and the extent to which distinct attributes become mixed. (scikit-learn.org)
In computer vision, designed descriptors summarize visual properties rather than treating every pixel as an independent attribute. The histogram of oriented gradients descriptor computes image gradients, groups their directions into local histograms, and normalizes groups of neighboring regions. The resulting vector describes local appearance through directional intensity changes. Normalization reduces sensitivity to some illumination variations, but the representation still depends on choices such as region size and orientation-bin count. (scikit-image.org)
Statistical extraction
Principal component analysis (PCA) constructs orthogonal directions that successively explain the greatest possible variance in centered numerical data. Keeping only some components produces a lower-dimensional representation. If , the projection can be written
where is the fitted mean and the columns of are the retained component directions. PCA can be computed using singular value decomposition. (sklearn.org)
PCA is an unsupervised learning method: its ordinary formulation uses the input measurements without target labels. Consequently, preserving high-variance directions is not the same objective as preserving the information most useful for a particular prediction task. Feature scaling also matters because variables measured on different numerical scales can influence the component directions differently. Centering alone does not standardize their scales. PCA components are combinations of original attributes, so a compact representation may be less directly interpretable than the original measurements. (sklearn.org)
Learned neural features
In deep learning, intermediate network activations serve as learned features. A convolutional neural network can supply image representations to a final classifier, with the feature-producing layers trained together with that classifier. Under supervised learning, the representation is shaped by the labeled task and its loss function, rather than by an independently specified descriptor alone. (docs.pytorch.org)
An autoencoder instead learns an encoder and decoder that attempt to reconstruct the input. Its internal code can function as an extracted representation. A restricted code size or other training constraint discourages trivial copying and can encourage useful structure. Nevertheless, successful reconstruction does not by itself establish that the code is optimal for a separate classification task. (deeplearningbook.org)
In transfer learning, a pretrained network may be used as a fixed extractor: its parameters remain frozen while a new prediction layer is trained on its outputs. This differs from fine-tuning, which updates some or all pretrained parameters for the new task. Fixed extraction separates representation reuse from learning the new predictor, whereas fine-tuning allows the representation itself to change. (docs.pytorch.org)
Evaluation and pipeline integrity
Extraction is part of the fitted analytical system, not merely a preliminary formatting operation. In ordinary inductive evaluation, learned quantities—including vocabularies, term weights, means, and projection directions—are estimated from training data. The fitted transformation is then applied unchanged to the test set. Estimating it from held-out observations can introduce data leakage and produce overly optimistic performance estimates. (scikit-learn.org)
During cross-validation, data-dependent extraction is fitted separately within each training fold. Evaluation consequently measures the combination of extractor and downstream model. At deployment, the system must preserve the same transformation, coordinate meanings, and preprocessing conventions used during training; otherwise, apparently valid numerical inputs may represent a different feature space. Pipelines make these dependencies explicit and help maintain consistency between fitting and subsequent use. (scikit-learn.org)