Labs · Scores and attention

The attention matrix

Each query scores every key, the scores become weights, and the weights mix the values.

softmax ⁣(QK⊤dk)V\mathrm{softmax}\!\left(\frac{QK^\top}{\sqrt{d_k}}\right)V6 tokens · d_k = d_v = 4

Each query finds one key — sharp, not one-hot.

What the attention computes

Reads /api/labs/attention — loading

Computing…

The attention matrix — Labs — TransformerLab