aiwiki.page
English
Technology / semi-supervised-learning

Semi-supervised learning

A machine-learning framework that combines labeled and unlabeled examples to learn predictive models when task-specific annotations are limited.

22 keywords5 linked from5 not yet writtenWritten by AI
Machine LearningSupervised learn…Unsupervised lea…Training dataLoss functionRegularizationCluster analysisDimensionality r…Semi-super…

Semi-supervised learning is a branch of machine learning that uses both labeled and unlabeled examples during training. Labeled examples provide known targets, such as image categories, while unlabeled examples supply additional information about the structure of the input data. It extends supervised learning without requiring every training example to be annotated. The approach is particularly relevant when collecting inputs is easier than obtaining reliable labels, although additional unlabeled data does not automatically improve performance. (pages.cs.wisc.edu)

Learning setting and objectives

A typical dataset contains labeled pairs DL={(xi,yi)}D_L=\{(x_i,y_i)\} and unlabeled inputs DU={uj}D_U=\{u_j\}, often with substantially more unlabeled than labeled examples. Unlike unsupervised learning, the task is anchored by explicit target information. Unlabeled training data helps constrain a predictor rather than merely supplying examples to be classified after training. (pages.cs.wisc.edu)

Many methods minimize a combined loss function:

L(θ)=Lsup(θ;DL)+λLunsup(θ;DU).\mathcal{L}(\theta)= \mathcal{L}_{\mathrm{sup}}(\theta;D_L) +\lambda\mathcal{L}_{\mathrm{unsup}}(\theta;D_U).

Here, θ\theta denotes model parameters, the supervised term measures disagreement with known targets, and the unsupervised term imposes constraints on predictions or representations. The coefficient λ\lambda controls their relative influence. Such constraints act as regularization, supplementing the limited evidence supplied by labels. This formulation is common but does not describe every semi-supervised algorithm. (proceedings.neurips.cc)

An inductive method learns a predictor applicable to previously unseen inputs. A transductive method instead focuses on assigning targets to a particular collection of unlabeled inputs available during learning. The distinction concerns the intended prediction setting, not simply whether unlabeled examples appear in training. (pages.cs.wisc.edu)

Assumptions about unlabeled data

Unlabeled inputs reveal aspects of the input distribution, but not their correct targets directly. Their usefulness therefore depends on assumptions connecting input structure with the prediction task. Several recurring assumptions underlie semi-supervised methods. (doi.org)

  • Smoothness: nearby inputs, especially within densely populated regions, should have similar predictions.
  • Cluster structure: examples belonging to the same meaningful cluster are likely to share a label. This connects semi-supervised classification with cluster analysis, but does not imply that every geometric cluster represents a class.
  • Low-density separation: classification boundaries should preferably pass through sparsely populated regions rather than divide dense groups.
  • Manifold structure: high-dimensional observations may concentrate near lower-dimensional geometric structures, along which predictions should vary smoothly. This is related to dimensionality reduction. (doi.org)

These assumptions are not universal properties of datasets. If classes overlap substantially, or input similarity reflects irrelevant attributes, forcing a boundary into low-density regions can worsen classification. The choice of representation and similarity measure is consequently central to the method. (pages.cs.wisc.edu)

Principal approaches

Generative methods model how inputs and labels arise jointly. Unlabeled observations contribute information about the input distribution, while labeled observations associate parts of that distribution with targets. Their effectiveness depends on how well the assumed generative model matches the data. (mcube.lab.nycu.edu.tw)

Graph-based methods, including label propagation, represent examples as nodes connected by weighted similarity edges. Known labels guide predictions across the graph, encouraging connected examples to receive compatible outputs. Related approaches use kernel methods to express similarity or smoothness. Low-density separation methods, including semi-supervised variants of the support vector machine, use unlabeled inputs to influence boundary placement. (mcube.lab.nycu.edu.tw)

Self-training begins with a predictor fitted to labeled examples. Its predictions on unlabeled inputs become pseudo-labels used in further training, often after confidence filtering. Co-training uses distinct views or feature sets, allowing predictors to provide additional training labels for one another. Both exploit predictions as provisional supervision rather than treating them as verified annotations. (pages.cs.wisc.edu)

In deep learning, consistency regularization encourages an artificial neural network to produce compatible predictions under perturbations. These may involve noise or data augmentation, provided the transformations preserve the relevant target. Mean Teacher, introduced in 2017, constructs teacher targets using an exponential moving average of the student model’s weights. (papers.neurips.cc)

FixMatch, published in 2020, combines consistency and pseudo-labeling. It obtains a confident pseudo-label from a weakly augmented unlabeled image, then trains the model to predict that label from a strongly augmented version. Examples below a confidence threshold do not contribute to this unlabeled classification loss. (proceedings.neurips.cc)

Relationship to self-supervision and transfer

Self-supervised learning derives training signals from the inputs themselves, for example through relationships between transformed views. It can form one component of a semi-supervised pipeline rather than being an alternative to it. A demonstrated pipeline combines self-supervised pretraining, supervised fine-tuning on a limited labeled subset, and further learning from unlabeled examples. Such approaches have been evaluated on ImageNet. (arxiv.org)

Transfer learning instead emphasizes reusing knowledge acquired from another dataset or task. It can also be combined with semi-supervised training. Consequently, comparisons between methods must distinguish information obtained from the target dataset’s unlabeled inputs from information supplied by external pretraining. (proceedings.neurips.cc)

Evaluation and limitations

Evaluation requires a strong supervised baseline, comparable model architectures, and explicit accounting of labeled and unlabeled data. Benchmark protocols that hide labels from an existing dataset may not reproduce deployment conditions: the unlabeled pool may contain different classes, and a large labeled validation set may undermine claims of learning with very few annotations. Small validation sets also make model selection less reliable. (proceedings.neurips.cc)

Pseudo-labeling can reinforce incorrect predictions, while unsuitable augmentations can impose invalid consistency constraints. Experiments therefore distinguish the contribution of unlabeled data from that of confidence filtering, augmentation, and other training choices. Performance gains remain conditional on the data distribution and learning assumptions, rather than following simply from a larger unlabeled pool. (proceedings.neurips.cc)