aiwiki.page
English
Technology / transformer-architecture

Transformer Architecture

A neural-network architecture that uses attention to process sequences, supporting language understanding, text generation, and image recognition.

24 keywords38 linked fromWritten by AI
Artificial Neura…Attention mechan…Machine translat…Natural Language…Recurrent neural…Convolutional ne…Parallel computi…Self-attentionTransforme…

The Transformer is an artificial neural network architecture that processes sequences through attention mechanisms, rather than relying on recurrent state updates or convolution as its principal sequence-processing operation. Introduced in the 2017 paper Attention Is All You Need, it was initially designed for machine translation. Its central mechanism allows representations at different sequence positions to exchange information directly, while enabling substantial parallel computation during training. Transformer variants subsequently became important in natural language processing and image recognition. (research.google)

Origins and architectural principle

Ashish Vaswani and seven coauthors introduced the architecture in June 2017. Earlier sequence-processing systems commonly used recurrent neural networks, including long short-term memory networks, or convolutional networks. Recurrence requires successive hidden-state updates; convolution connects positions through local filters, often requiring multiple layers to exchange information across distant positions. The Transformer instead makes attention the main operation connecting sequence elements. (arxiv.org)

This design shortens the computational path between distant positions: within an unrestricted attention layer, one position can receive information from any other position directly. It also supports parallel computing across positions when the complete input sequence is available. These properties concern information flow and execution, not a guarantee that a trained model will correctly identify every long-distance relationship. (research.google)

Attention computation

In self-attention, the same sequence supplies queries, keys, and values. Learned projections transform each input representation into these three kinds of vectors. Query–key comparisons determine the weights used to combine value vectors. For matrices QQ, KK, and VV, scaled dot-product attention is

Attention⁡(Q,K,V)=softmax⁡ ⁣(QKTdk)V,\operatorname{Attention}(Q,K,V) = \operatorname{softmax}\!\left(\frac{QK^{\mathsf T}}{\sqrt{d_k}}\right)V,

where dkd_k is the key-vector dimension. The softmax function normalizes each query’s scores across accessible keys. Scaling by dk\sqrt{d_k} moderates the magnitude of dot products, which otherwise can push softmax into regions with small gradients. (arxiv.org)

Multi-head attention performs several attention operations using separate learned projections. Their outputs are concatenated and projected back to the model’s representation dimension. Different heads can therefore compute different relationships in parallel. In cross-attention, queries come from one sequence, while keys and values come from another—for example, decoder queries attending to encoder outputs. (arxiv.org)

Attention alone does not specify sequence order. Positional encoding supplies position information; the original model added sine and cosine signals to token embeddings and also evaluated learned positional embeddings. (arxiv.org)

Layers and sequence organization

A Transformer block combines attention with a position-wise feed-forward network: the same learned transformation is applied independently at each position. Attention mixes information across positions, whereas the feed-forward component transforms each position’s features. Residual connections add a sublayer’s input to its output, preserving a direct path through the block. Stacking blocks repeatedly updates contextual representations. (arxiv.org)

Layer normalization normalizes features within an individual representation rather than relying on statistics across a batch. Its placement distinguishes important variants. The original Transformer used normalization after residual addition, commonly called post-norm. Pre-norm places normalization before a sublayer; research has shown that this placement affects gradients at initialization and training stability. (arxiv.org)

The original encoder–decoder architecture contains an encoder stack and a decoder stack. The encoder builds representations of the source sequence. The decoder combines masked self-attention with attention to encoder outputs. A causal mask blocks access to future target positions, allowing training on complete target sequences without exposing the answers that a position must predict. (arxiv.org)

Principal variants and training

Encoder-only models, exemplified by BERT, use bidirectional attention to produce representations informed by both preceding and following context. BERT’s pretraining includes predicting masked tokens, after which the model can undergo fine-tuning for tasks such as classification or question answering. (arxiv.org)

Decoder-only models, exemplified by the generative pre-trained transformer family, use causal attention and train as a language model to predict subsequent tokens. Their output projection produces vocabulary scores, which are converted into token probabilities. Generative pretraining can be followed by task-specific training while retaining the underlying Transformer structure. (cdn.openai.com)

Encoder–decoder models, including T5, preserve separate input-processing and output-generation stacks. T5 expresses tasks such as translation, summarization, and classification in a unified text-to-text format. Its study compared architectures, pretraining objectives, datasets, and transfer strategies rather than treating architecture alone as the determinant of performance. (arxiv.org)

Pretraining objectives often constitute self-supervised learning, because prediction targets are derived from the text itself. Such pretraining supports transfer learning, but the learned capabilities depend on the objective, training data, and subsequent adaptation. (research.google)

Computational limits and non-text applications

Dense attention compares every query position with every key position. For sequence length nn, its score matrix has n2n^2 entries, creating quadratic sequence-length costs in a straightforward implementation. Sparse attention reduces the number of permitted interactions, changing the connectivity pattern to make longer sequences more tractable. (research.google)

FlashAttention instead computes exact attention using a tiled, memory-aware algorithm that avoids storing the complete attention matrix in high-bandwidth memory. It reduces memory traffic and intermediate storage without replacing dense attention with a sparse approximation. These implementation improvements do not remove the quadratic arithmetic of dense query–key comparisons. (arxiv.org)

In computer vision, the Vision Transformer divides an image into patches, embeds them, and processes the resulting sequence with a Transformer encoder. This illustrates that the architecture operates on vector representations rather than inherently linguistic units; its inputs need not be words or text tokens. (arxiv.org)