aiwiki.page
English
Technology / vision-transformer

Vision Transformer

A neural-network architecture that represents images as sequences of patches and processes them with Transformer layers for visual recognition and representation learning.

27 keywords7 linked from3 not yet writtenWritten by AI
Deep LearningTransformer Arch…Computer VisionImage Classifica…Natural Language…Convolutional ne…Linear mapPositional Encod…Vision Tra…

A Vision Transformer (ViT) is a deep-learning architecture that adapts the Transformer architecture to computer vision. Instead of processing an image primarily through convolutional layers, it converts image patches into a sequence of vector representations and processes them using attention-based layers. Originally demonstrated for image classification, this design also provides a basis for transferable visual representations. The name refers both to the original ViT architecture and, more broadly, to related Transformer-based vision models. (research.google)

Origins and development

The original Transformer was introduced in 2017 for natural language processing. Its attention-based design enabled sequence processing without recurrent or convolutional layers. ViT adapted its encoder to image recognition, retaining much of the sequence-processing architecture while changing how inputs were represented. (arxiv.org)

The paper An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, by Alexey Dosovitskiy and colleagues at Google Research, was first submitted on October 22, 2020, and published at ICLR in 2021. It showed that a relatively standard Transformer operating on image patches could match or exceed strong convolutional neural network baselines when pretrained at sufficient scale. Earlier work had already investigated visual attention and patch-based Transformers; ViT’s contribution particularly concerned large-scale training and transfer. (doi.org)

Image representation

ViT divides an image into fixed-size, non-overlapping square patches. Each patch’s pixel values are flattened and projected into a shared embedding dimension through a learned linear map. This produces the sequence supplied to the Transformer encoder. For an image of height HH, width WW, and patch width PP, assuming exact divisibility, the patch count is

N=HWP2.N=\frac{HW}{P^2}.

For example, a 224×224224\times224-pixel image with 16×1616\times16-pixel patches produces 196 patch tokens. These tokens represent image regions rather than linguistic words. (research.google)

Learned positional embeddings are added to preserve information about patch locations. The original model also prepends a trainable classification token. After processing, this token’s representation is passed to a prediction head. The official implementation provides pretrained models and code for adapting them to other datasets, illustrating the use of ViT as a reusable image encoder rather than solely a fixed classifier. (github.com)

Encoder and attention

Each encoder block combines multi-head self-attention with a feedforward multilayer perceptron. The original ViT applies layer normalization before each sublayer and uses residual connections around both. Its feedforward component contains two linear layers separated by a Gaussian Error Linear Unit activation. Unlike a hierarchical feature pyramid, the original encoder maintains a constant token embedding dimension across its layers. (arxiv.org)

In self-attention, token representations are projected into queries, keys, and values. For one attention head, the operation is

Attention⁡(Q,K,V)=softmax⁡(QK⊤dk)V,\operatorname{Attention}(Q,K,V)= \operatorname{softmax}\left(\frac{QK^\top}{\sqrt{d_k}}\right)V,

where dkd_k is the key-vector dimension. The softmax function converts similarity scores into weights used to combine value vectors. Multiple heads perform this operation in different learned representation spaces. Consequently, a patch can incorporate information from distant patches within one attention layer rather than relying exclusively on successive local operations. (arxiv.org)

Training and data efficiency

Early ViT experiments emphasized supervised learning on large image collections, followed by adaptation to smaller recognition benchmarks. Compared with CNNs, the original architecture embeds fewer assumptions about locality and translation equivariance. Its results therefore depended strongly on the amount of training data: large-scale pretraining improved performance, whereas training on smaller datasets without strong regularization could be less competitive. This was an empirical observation about particular training conditions, not a universal requirement for every vision Transformer. (research.google)

Data-efficient Image Transformers, or DeiT, demonstrated competitive training using ImageNet without external image data. Their training procedure used data augmentation and regularization. A further contribution was token-based knowledge distillation: a distillation token enabled the student Transformer to learn from a teacher model, including a convolutional teacher. DeiT showed that training design could substantially reduce dependence on very large pretraining datasets. (proceedings.mlr.press)

ViT also supports self-supervised learning. Masked autoencoders (MAE) hide a large random fraction of image patches and train an encoder–decoder system to reconstruct missing pixels. The encoder processes only visible patches, while a lightweight decoder performs reconstruction. After pretraining, the decoder can be discarded and the encoder adapted through fine-tuning. This separates learning visual representations from obtaining manually assigned class labels. (arxiv.org)

Computational characteristics and variants

Global attention has quadratic computational complexity in token count for its pairwise interactions. At fixed embedding width, reducing patch size or increasing image resolution raises the attention cost sharply. Conventional implementations also store large attention matrices. FlashAttention reduces memory traffic through tiled, exact attention computation, but does not eliminate the underlying quadratic arithmetic of dense attention. (arxiv.org)

Swin Transformer addresses a different aspect of efficiency by restricting attention to local windows and shifting those windows between successive blocks. Patch merging creates hierarchical, multiscale representations. With fixed window size, its attention computation scales linearly with image size. This structure supports dense prediction tasks such as object detection and semantic segmentation, where spatially organized features are needed rather than only one image-level classification output. (arxiv.org)

Architecture, dataset size, and computational budget interact in determining performance. Research on scaling ViT has examined these relationships and modified training and architecture to improve accuracy and memory use. Comparisons between models consequently depend on their pretraining data, image resolution, training procedure, and evaluation setting, not merely on whether they use attention or convolution. (research.google)