Transfer learning is an approach in machine learning that uses knowledge acquired from a source task or domain to improve learning in a target task or domain. The transferred knowledge may consist of model parameters, learned features, selected examples, or relationships among data. Unlike learning entirely from target training data, it draws on information obtained elsewhere. It is particularly useful when target data are limited or expensive to label, although successful transfer depends on the relationship between source and target problems. (cse.cuhk.edu.hk)
Domains, tasks, and related concepts
A common formalization distinguishes a domain, consisting of a feature space and a marginal probability distribution over inputs, from a task, consisting of an output space and a predictive function. Transfer learning concerns situations in which the source and target domains, tasks, or both differ. Thus, transferring a classifier between different image collections and adapting a pretrained model to a different prediction objective are distinct forms of transfer. (cse.cuhk.edu.hk)
Domain adaptation generally addresses changes in data distribution while retaining a related prediction task. Multi-task learning emphasizes learning several tasks jointly, whereas transfer learning typically emphasizes improvement on a designated target. Transfer can operate in supervised, unsupervised, and other learning settings; it is not restricted to neural networks or labeled source data. (cse.cuhk.edu.hk)
What is transferred
In deep learning, transfer commonly relies on representation learning: a source model encodes inputs into features that can support another prediction problem. A pretrained network therefore provides both a computational mapping and parameters shaped by earlier experience. Reusing these parameters can supply a more useful starting point than random initialization, rather than simply importing the source model’s final predictions. (arxiv.org)
Experiments with convolutional neural networks trained on natural images found that early layers often capture relatively general patterns, such as edges and color contrasts, while later layers become more specialized. This is an empirical tendency, not a universal rule for all architectures and datasets. Transferability also depends on interactions among layers: separating co-adapted components can create optimization difficulties even when individual features appear useful. (arxiv.org)
Distribution differences may require more than copying weights. Deep Adaptation Networks, for example, combine task learning with an objective that reduces discrepancies between source and target representations. Such methods seek features that remain useful across domains rather than assuming that source-trained features already match target inputs. (arxiv.org)
Feature extraction and fine-tuning
Two common neural-network workflows differ in how much of the pretrained model changes:
- Fixed feature extraction: the pretrained backbone remains frozen, and a new prediction head is trained on its outputs.
- Fine-tuning: some or all pretrained parameters are updated using target data, allowing the representation itself to change. (keras.io)
A typical workflow replaces the original output layers, trains the new head, and optionally unfreezes part of the backbone. A relatively small learning rate during fine-tuning can limit abrupt changes to previously learned features. Freezing parameters and setting a model’s inference behavior are separate operations; layers such as batch normalization can maintain statistics that require special handling during adaptation. (keras.io)
These approaches offer different trade-offs. A frozen backbone restricts target adaptation but reduces the number of trainable parameters. Fine-tuning provides greater flexibility, while increasing the possibility of overfitting on small target datasets or disrupting useful pretrained features. Neither workflow is inherently superior for every source–target pairing. (keras.io)
Applications and development
In computer vision, models pretrained on ImageNet can provide reusable image representations. A classification backbone may be retained while its original output layer is replaced for another image-classification problem. Research published in 2014 systematically examined how layer depth, task differences, and subsequent fine-tuning influence this transfer. (keras.io)
In natural language processing, pretraining on unlabeled text followed by task-specific adaptation became an important transfer mechanism. BERT, presented at NAACL in 2019, demonstrated that pretrained bidirectional Transformer representations could be fine-tuned with a small additional output component for tasks including question answering and language inference. Its pretraining objectives are examples of self-supervised learning, in which training signals are constructed from the input data itself. (aclanthology.org)
Parameter-efficient fine-tuning changes only a limited subset of parameters or adds compact trainable components. Adapter-based transfer, introduced in a 2019 study, inserts small modules into a pretrained network while retaining shared backbone parameters. This reduces task-specific storage compared with maintaining a fully fine-tuned model for every task. Later research analyzed connections among several parameter-efficient approaches and their design choices. (arxiv.org)
Limitations and evaluation
Transfer does not guarantee better generalization. Negative transfer occurs when using source knowledge worsens target learning relative to an appropriate target-only baseline. Feature-transfer experiments show that increasing differences between source and target tasks can reduce transferability, although even relatively distant source features may sometimes outperform random features. (cse.cuhk.edu.hk)
Evaluation distinguishes model selection from final performance measurement. Cross-validation or a validation set can support selection of adaptation settings, while an independent test set estimates performance on unseen examples. Data leakage, including fitting preprocessing steps using test information, can produce misleadingly optimistic results. (scikit-learn.org)
Privacy is another limitation. Experiments have shown that access to a fine-tuned model can sometimes reveal whether particular examples belonged to its pretraining dataset. Adapting a model to a new task therefore does not necessarily eliminate information retained about source data. (arxiv.org)