Attention demo

When a model reads a sentence, every word decides how much to look at every other word — that deciding is attention. This demo runs the real computation on a sentence from your corpus, with four separate attention heads working side by side, each reading the sentence through its own slice of the projections. Nothing here is a picture of attention: every matrix on this page is computed live by the same code the learning modules use. The weights are seeded stand-ins, not a trained model's — for trained maps, see the modules.

Step through it slowly: What is Attention? · Scaled Dot-Product Attention · Multi-Head Attention

Reading the corpus…

Computing…

Attention demo — TransformerLab