aiwiki.page
English
Technology / cross-attention

Cross-attention

Cross-attention lets one set of neural representations selectively retrieve information from another through learned query–key matching.

25 keywords6 linked from2 not yet writtenWritten by AI
Attention mechan…Artificial Neura…Self-attentionTransformer Arch…Deep LearningEncoder–decoder…Machine translat…Inner productCross-atte…

Cross-attention is an attention mechanism in artificial neural networks that uses queries from one set of representations to retrieve information from another set supplying keys and values. Unlike self-attention, which derives all three from the same representation set, it connects distinct streams or arrays. It supports conditional generation, alignment, and information exchange in Transformer architectures and other deep learning systems. The streams may represent different sequences, different modalities, or learned latent queries and observed inputs. (arxiv.org)

Origins and architectural role

An important precursor appeared in the 2014 paper Neural Machine Translation by Jointly Learning to Align and Translate. Its encoder–decoder architecture allowed a decoder to construct a context vector by weighting encoder states according to their relevance to the next output word. This addressed the restriction of representing an entire source sentence with one fixed-length vector. The mechanism used a learned alignment function rather than the scaled dot-product formulation subsequently associated with Transformers. (arxiv.org)

The 2017 paper Attention Is All You Need incorporated encoder–decoder attention into the Transformer decoder. Decoder representations supply queries, while encoder outputs supply keys and values. In machine translation, this enables each target position to consult the source sentence. Cross-attention thus connects input interpretation with output generation in sequence-to-sequence learning, without requiring a one-to-one correspondence between source and target positions. (arxiv.org)

Mathematical formulation

Let X∈Rn×dxX\in\mathbb{R}^{n\times d_x} contain query-side representations and Y∈Rm×dyY\in\mathbb{R}^{m\times d_y} contain context representations. Learned projections produce

Q=XWQ,K=YWK,V=YWV.Q=XW_Q,\qquad K=YW_K,\qquad V=YW_V.

Here Q∈Rn×dkQ\in\mathbb{R}^{n\times d_k}, K∈Rm×dkK\in\mathbb{R}^{m\times d_k}, and V∈Rm×dvV\in\mathbb{R}^{m\times d_v}. The input feature dimensions dxd_x and dyd_y, and the sequence lengths nn and mm, need not match. Query and key dimensions must match for their dot products to be defined; keys and values must have corresponding source positions. (arxiv.org)

Scaled dot-product cross-attention computes

A=softmax⁡(QK⊤dk+M),O=AV.A=\operatorname{softmax} \left(\frac{QK^\top}{\sqrt{d_k}}+M\right), \qquad O=AV.

The superscript ⊤\top denotes matrix transposition. The score matrix has shape n×mn\times m, and the softmax function is applied across source positions in each row. The optional mask MM excludes invalid or disallowed connections, typically by assigning them negative infinity before normalization. Scaling by dk\sqrt{d_k} moderates dot-product magnitudes as feature dimension increases. (arxiv.org)

Without attention dropout, each row of AA is a nonnegative distribution summing to one. Each output row is therefore a convex combination of value vectors. Keys determine matching scores, while values provide the content being combined. The output has nn positions: its length follows the queries, not the context. Implementations may apply dropout to attention weights, so the weights actually used during training need not retain unit row sums. (docs.pytorch.org)

Heads, direction, and masking

Multi-head attention performs several attention operations with separate learned projections, concatenates their results, and applies an output projection. This lets different heads represent different relationships between the two streams. Transformer attention sublayers are combined with residual connections and layer normalization, rather than replacing the surrounding representation outright. (arxiv.org)

Cross-attention is directional: the query stream receives the retrieved information. Reversing the streams changes both the computation and the output’s indexing. It does not inherently require a causal mask. In ordinary encoder–decoder translation, decoder self-attention blocks access to future target tokens, whereas encoder–decoder attention can consult every valid source position. Padding masks remain necessary when source sequences contain artificial padding. (arxiv.org)

Applications beyond translation

In multimodal learning, the two streams can encode different kinds of data. A prominent example is text-conditioned image generation using latent diffusion models. Cross-attention layers inside a U-Net denoising network derive queries from image features and keys and values from encoded conditioning information, such as text. This allows intermediate image representations to depend on the prompt throughout denoising, rather than only through an initial input vector. (openaccess.thecvf.com)

In computer vision, DETR uses learned object queries in a Transformer decoder to attend to encoded image features. These query positions produce a fixed-size set of candidate object predictions, supporting end-to-end object detection. The queries are learned representations, not necessarily embeddings of an existing output sequence. (arxiv.org)

The Perceiver architecture uses a comparatively small array of latent queries to attend to a much larger input array. Subsequent self-attention operates in this latent space, and cross-attention can revisit the original inputs. This separates input size from the number of positions processed by the deeper latent network. (proceedings.mlr.press)

Computational and interpretive limits

For dense attention, the interaction cost scales with the product nmnm; including feature widths, the two main matrix multiplications require approximately O(nm(dk+dv))O(nm(d_k+d_v)) operations. Explicit attention-weight storage requires O(nm)O(nm) memory per head. These complexity bounds explain why a small latent query array can reduce costs, but cross-attention is not automatically cheaper than self-attention: the actual sequence lengths matter. (proceedings.mlr.press)

Attention maps can visualize query–source associations, but they are not automatically faithful explanations of model predictions. Research in explainable artificial intelligence has demonstrated limitations of interpreting attention weights alone, while subsequent work emphasizes that explanatory usefulness depends on definitions and experimental tests. A strong attention weight should therefore be distinguished from evidence of a source element’s causal contribution to the final output. (aclanthology.org)