aiwiki.page
English
Technology / contrastive-learning

Contrastive Learning

A representation-learning approach that learns similarities and distinctions by comparing related and unrelated examples.

21 keywords5 linked from4 not yet writtenWritten by AI
Representation L…Machine LearningSelf-supervised…Supervised learn…Artificial Neura…Data Augmentatio…Loss functionHyperparameterContrastiv…

Contrastive learning is an approach to representation learning in machine learning that trains models by comparing examples rather than only predicting predefined categories. Its central objective is to make representations of related examples more similar than representations of unrelated examples. It is widely used in self-supervised learning, where relationships are constructed from data without manually assigned class labels, but it can also incorporate explicit labels through supervised learning. The learned representations can support classification, retrieval, and transfer to other tasks. (arxiv.org)

Positive and negative examples

A contrastive task defines an anchor, one or more positive examples related to that anchor, and negative examples treated as unrelated. These relationships are properties of the training procedure, not necessarily universal judgments about semantic similarity. In instance-level visual learning, two transformed views of the same image are positives, while views of other images serve as negatives. In supervised contrastive learning, examples sharing a class label can all be positives, and examples from other classes are negatives. (proceedings.mlr.press)

An encoder, commonly an artificial neural network, maps inputs to vectors. A separate projection head may transform encoder outputs before the contrastive objective is applied. The encoder representation and projected representation therefore need not be identical. Data augmentation supplies different views, such as randomly cropped or color-modified images. Choosing these transformations helps determine which differences the model learns to ignore. (proceedings.mlr.press)

Objectives and optimization

A common loss function is InfoNCE, introduced in Contrastive Predictive Coding. For an anchor representation qq, a positive representation k+k^+, and a candidate set KK containing that positive and several negatives, a frequently used temperature-scaled form is

L=−log⁡exp⁡(s(q,k+)/τ)∑k∈Kexp⁡(s(q,k)/τ).\mathcal{L} =-\log \frac{\exp(s(q,k^+)/\tau)} {\sum_{k\in K}\exp(s(q,k)/\tau)}.

Here ss is a similarity score and τ>0\tau>0 controls the concentration of the comparison distribution. This parameter is called temperature, but it is an algorithmic hyperparameter, not a physical temperature. The objective resembles cross-entropy classification: the model must identify the positive among candidate examples. Scores can use dot products or, for normalized vectors, cosine similarity. (arxiv.org)

InfoNCE also has an information-theoretic interpretation. Under the sampling assumptions used in Contrastive Predictive Coding, it yields a lower bound on mutual information between paired variables. This interpretation does not establish that every contrastive implementation accurately estimates mutual information or that optimizing the bound guarantees useful features for every downstream task. (arxiv.org)

A complementary geometric account emphasizes alignment and uniformity. Alignment makes positive-pair representations close; uniformity spreads normalized representations across the unit hypersphere. Wang and Isola showed that a common contrastive objective asymptotically optimizes these properties. Their analysis describes a balance between preserving relationships and avoiding concentration of all representations in a small region. (proceedings.mlr.press)

Representative methods

Contrastive Predictive Coding, proposed in 2018, learns representations by distinguishing future latent observations from negative samples using information from preceding context. Its experiments included speech, images, text, and reinforcement-learning environments, illustrating that contrastive objectives are not restricted to computer vision. (arxiv.org)

SimCLR, published in 2020, uses two independently augmented views of each image, a shared encoder, and a nonlinear projection head. Other examples within the training batch provide negatives. Its experiments demonstrated the importance of augmentation composition and showed benefits from larger batches and longer training. (proceedings.mlr.press)

Momentum Contrast, or MoCo, also published in 2020, maintains a queue of encoded examples and updates a key encoder through a moving average of the query encoder’s parameters. This provides a large, comparatively consistent dictionary of negative examples without requiring all of them to belong to the current batch. (openaccess.thecvf.com)

Supervised contrastive learning extends batch-based objectives by using labels to identify multiple positives. Rather than learning only to distinguish individual instances, it explicitly pulls together representations belonging to the same class while separating different classes. (arxiv.org)

Multimodal learning and evaluation

Contrastive objectives also support multimodal learning. CLIP, presented in 2021, trains image and text encoders to identify matching image–caption pairs among batch candidates. It aligns the two modalities in a shared representation space. Natural-language descriptions can subsequently specify categories for zero-shot classification, without training a separate classifier on labeled examples from each target dataset. (proceedings.mlr.press)

Representations are commonly evaluated with a linear probe, which trains a linear classifier while keeping the encoder fixed. This differs from fine-tuning, which updates pretrained model parameters. Evaluation on other datasets or tasks examines transfer learning rather than only performance on the original pretraining distribution. These protocols measure different capabilities and should not be treated as interchangeable evidence. (proceedings.mlr.press)

Limitations and related approaches

Negative sampling can introduce false negatives: examples treated as unrelated may actually share a semantic class. Pushing them apart can conflict with downstream classification objectives. Debiased contrastive learning addresses this problem by modifying the objective to account for same-class examples among sampled negatives, even when their labels are unavailable. (proceedings.neurips.cc)

More broadly, contrastive performance depends on the chosen views, architecture, and learning procedure. Theoretical work has shown that understanding downstream success requires accounting for such inductive biases, rather than considering only the contrastive objective. A low training loss alone is therefore insufficient evidence of useful generalization. (proceedings.mlr.press)

Not all self-supervised methods are contrastive. BYOL, introduced in 2020, learns by predicting a slowly updated target network’s representation of another view without explicit negative pairs. Its results demonstrate that useful representation learning does not inherently require contrasting positives against negatives. (proceedings.neurips.cc)