Labs · Scores and attention
Heads split and mixed
One attention split into heads by slicing the projections, each head's map on its own, the contexts concatenated and mixed back by W_O.
6 tokens · d_model = 8 · 2 heads of d_k = 4
Head 0 attends one token back, head 1 to the matching token — weights constructed to show the split.
Each head's map, the concatenation, the mix
Reads /api/labs/heads — loading
Computing…