Skip to main content

Attention Mechanism

Suppose two pieces of information are represented by these vectors:

v1=[2,0],v2=[0,4].\mathbf{v}_1=[2,0],\quad \mathbf{v}_2=[0,4].

Taking half of each and adding gives

12[2,0]+12[0,4]=[1,2].\tfrac12[2,0]+\tfrac12[0,4]=[1,2].

But how much should we take from each? Attention computes these weights from the content rather than fixing them at one half.

This page explains the Q/K/V weighted-mixing operation, masks, and normalization. For the origin of its input vectors, start with tokenization and text representations; for assembling a complete network, continue to Transformer. Multimodal models explains how image patches connect to text.

How Much Weight Should Each Vector Receive?​

The vectors being mixed are called values. To choose their weights, introduce a query, the vector used to retrieve information.

Each value has a corresponding key, which is compared with the query to compute a matching score. In this example, take

q=[1,0],k1=[1,0],k2=[0,1].\mathbf{q}=[1,0],\quad \mathbf{k}_1=[1,0],\quad \mathbf{k}_2=[0,1].

First compute weights using only the query and keys. The dot products are q⊤k1=1\mathbf q^\top\mathbf k_1=1 and q⊤k2=0\mathbf q^\top\mathbf k_2=0. Each key has two components. Divide by 2\sqrt2, the square root of that dimension, to get scores [1/2,0]≈[0.707,0][1/\sqrt2,0]\approx[0.707,0].

Softmax turns these scores into nonnegative weights that sum to one: it exponentiates the scores and divides by their sum.

α1=e1/2e1/2+1≈0.67,α2=1e1/2+1≈0.33.\alpha_1=\frac{e^{1/\sqrt2}}{e^{1/\sqrt2}+1}\approx0.67, \qquad \alpha_2=\frac{1}{e^{1/\sqrt2}+1}\approx0.33.

Now mix the values, keeping each weight paired with the value at the same position. Using rounded weights, the output is

0.67[2,0]+0.33[0,4]≈[1.34,1.32].0.67[2,0]+0.33[0,4]\approx[1.34,1.32].

The first key matches better, but the output still contains the second value. Temperature, score gaps, duplicate keys, and masks all change the mixture. “Attending to a position” is usually not a binary fact.

Change either component of q\mathbf q below; the keys and values stay fixed. At q=[0,0]\mathbf q=[0,0], both scores are zero, the weights are [1/2,1/2][1/2,1/2], and the output is [1,2][1,2]. At [0,1][0,1], the weights reverse and the output is approximately [0.66,2.68][0.66,2.68]. Scores determine the proportions; values determine what is mixed.

xyv₁ = [2, 0]v₂ = [0, 4]O

k1 = [1, 0]Score 0.707

Weight 67.0%

k2 = [0, 1]Score 0.000

Weight 33.0%

Output O: [1.340, 1.321]Each score is q · kᵢ / √2; softmax turns the two scores into weights that sum to 1. O = w₁v₁ + w₂v₂ is the weighted mixture of the two value vectors. Changing q moves O along the segment between v₁ and v₂.

For any number of key/value pairs, this same computation is

Attention⁡(q,K,V)=∑iαivi,αi=exp⁡a(q,ki)∑jexp⁡a(q,kj).\operatorname{Attention}(\mathbf{q},K,V)=\sum_i\alpha_i\mathbf{v}_i, \qquad \alpha_i=\frac{\exp a(\mathbf{q},\mathbf{k}_i)}{\sum_j\exp a(\mathbf{q},\mathbf{k}_j)}.

Query, key, and value are computational roles, not intrinsic meanings. They are usually obtained from inputs through different learned projections, and values need not equal keys. Computing weights from content and mixing values this way is a differentiable operation called content addressing.

Scores​

Additive attention uses a small learned network:

a(q,k)=v⊤tanh⁡(Wqq+Wkk).a(\mathbf{q},\mathbf{k})=\mathbf{v}^{\top}\tanh(W_q\mathbf{q}+W_k\mathbf{k}).

Scaled dot-product attention uses

a(q,k)=q⊤kdk.a(\mathbf{q},\mathbf{k})=\frac{\mathbf{q}^{\top}\mathbf{k}}{\sqrt{d_k}}.

If the query and key coordinates are mutually independent, zero-mean, and unit-variance, the unscaled dot product has variance dkd_k and standard deviation dk\sqrt{d_k}, which can saturate Softmax. The scaling is a design for this score, not a law for all attention.

Matrix Form and Masks​

For Q∈Rnq×dkQ\in\mathbb{R}^{n_q\times d_k}, K∈Rnk×dkK\in\mathbb{R}^{n_k\times d_k}, and V∈Rnk×dvV\in\mathbb{R}^{n_k\times d_v}:

A=softmax⁡ ⁣(QK⊤dk+M),O=AV.A=\operatorname{softmax}\!\left(\frac{QK^{\top}}{\sqrt{d_k}}+M\right), \qquad O=AV.

AA has shape nq×nkn_q\times n_k and OO has shape nq×dvn_q\times d_v. Masks act before Softmax:

  • Padding masks exclude placeholder keys.
  • Causal masks exclude future positions.
  • Structural masks restrict local, graph, or block connections.

With an additive −∞-\infty mask, a wholly masked row has only −∞-\infty logits and ordinary softmax has no defined probability distribution. Check the behavior of the chosen API and handle such rows explicitly; there is no universal return value. Changing a masked token should not change an unmasked output, provided the allowed Q/K/V representations are otherwise fixed.

Normalize Over Keys, Not Queries​

Softmax in AA acts independently across each row (the key axis). For additive masks, set Mij=0M_{ij}=0 for allowed pairs and −∞-\infty for forbidden pairs. In a causal sequence of three positions, row one can see key one; row two can see keys one and two; row three can see all three. With all allowed scores equal to zero, the rows are [1,0,0][1,0,0], [1/2,1/2,0][1/2,1/2,0], and [1/3,1/3,1/3][1/3,1/3,1/3]. The diagonal is allowed when the input at that position predicts the next token. This is the masking convention in the original Transformer.

A key-padding mask does not automatically remove padded query outputs; exclude those outputs from the loss or subsequent pooling. Boolean-mask conventions differ between APIs, so verify whether True means allowed or forbidden. The masked-token invariance test assumes Q/K/V at allowed positions are otherwise fixed: upstream operations must not already have mixed the masked content into them.

Self, Cross, and Multi-Head Attention​

  • Self-attention: Q/K/V come from the same set of representations, usually through different learned projections.
  • Cross-attention: Queries come from one side, while keys and values come from another.
  • Multi-head attention: Attention runs in parallel in several projected subspaces, then the outputs are concatenated and projected.

Multiple heads provide multiple computational channels, without guaranteeing that each maps to a stable human concept such as syntax or coreference. Heads can be redundant or change under reparameterization while preserving similar outputs.

Bahdanau-style attention first gave each RNN decoder step access to multiple encoder states to construct a different context; Luong compared global and local forms. The Transformer later made self-attention a primary layer operation. Attention did not begin with the Transformer.

Efficiency Designs Solve Different Problems​

Dense self-attention forms n×nn\times n scores; at fixed head dimensions, its naive time and intermediate memory grow quadratically with sequence length. The designs below target different bottlenecks.

DesignChangeDoes not automatically solve
FlashAttentiontiled, IO-aware exact attention without storing the full intermediate matrixarithmetic remains usually quadratic; kernel/hardware dependence
MQAquery heads share one K/V headcapacity or quality may decline; training suitability and compatibility still need verification
GQAgroups of query heads share a smaller set of K/V headsa cache/quality compromise, not a new scoring function
sliding-window/sparserestrict reachable positionsomitted positions need cross-layer paths
linear/kernel attentionreorder or approximate aggregationnormalization, numerics, and quality need not equal Softmax

MQA/GQA mainly reduce autoregressive KV-cache storage and memory traffic without reducing the query-head count. They change sharing, not the scoring function; their names alone do not establish quality, training suitability, or compatibility. Report query heads, K/V heads, head dimension, window, dtype, and kernel rather than only saying “efficient attention.” RNNs/SSMs may suit fixed-state streaming; convolution may use less data and computation when locality is strong. Attention is not the default answer for every relationship.

Interpretation Boundary​

English-to-French attention weights: brighter cells carry more weight; zone attends to Area despite the changed word order.Open full-size image

Read across one row: it shows the source-word weights used for one generated French word. White means weight 1 and black means 0. Around European Economic Area, the bright cells leave the main diagonal: zone attends to Area, then économique and européenne move back through the English phrase. This is recurrent encoder–decoder attention with additive scoring, not a Transformer self-attention head. The heatmap shows learned alignment; causal importance still requires the checks below.

Attention weights describe routing coefficients in a particular forward pass, but usually cannot establish causal importance on their own:

  • Different weights may yield similar outputs.
  • Values and later layers change the final influence.
  • Gradients, counterfactual replacements, and interventions may produce different rankings.
  • Neither a very sharp nor a very flat distribution automatically indicates good or bad behavior.

“Attention is not Explanation” demonstrates failures of direct interpretation; a competing analysis argues attention can still be informative under defined diagnostics and counterfactual tests. The defensible conclusion is that weights may be evidence, but should be combined with deletion, replacement, or other interventions and task validation. They cannot serve as a standalone causal explanation.

Connections to the next explanation​

The normalization is the same operation explained in Softmax Regression, but attention weights mix values rather than predict class labels. Transformer puts attention together with position information, residual connections, normalization and feed-forward layers. Attention variants and KV cache follows the same Q/K/V shapes into autoregressive inference and memory costs.

This Q/K/V account is biased toward NLP and Transformer practice. Vision, sets, graphs, and earlier statistical kernel formulations receive only boundary coverage.

Explore connectionsOpen network