Labs · Scores and attention

Heads split and mixed

One attention split into heads by slicing the projections, each head's map on its own, the contexts concatenated and mixed back by W_O.

MHA(X)=[head1 ∥ ⋯ ∥ headH] WO\mathrm{MHA}(X) = [\mathrm{head}_1 \,\|\, \cdots \,\|\, \mathrm{head}_H]\, W_O6 tokens · d_model = 8 · 2 heads of d_k = 4

Head 0 attends one token back, head 1 to the matching token — weights constructed to show the split.

Each head's map, the concatenation, the mix

Reads /api/labs/heads — loading

Computing…

Heads split and mixed — Labs — TransformerLab