Attention is simpler than you think - A hand-crafted superhero transformer TechAarvam's workshop materials include a hand-crafted transformer example that demonstrates the intuition behind the attention block. The notebook uses a simple vocabulary and hand-coded attention weights to show how attention extracts information for a feedforward network, which is trained to predict corrections to word attributes. Part of the TechAarvam workshop support files — Build Your Own Model https://www.techaarvam.com/workshops/build-your-own-model . This notebook presents a hand-crafted example. The goal is to understand the intuition behind the attention block in the transformer architecture. Attention and FFN are the two main components. FFN is an Artificial neural network ANN with 1 input, 1 hidden, and 1 output layer. So we start constructing the FFN's job by hand. Then we show how Attention extracts what the FFN can use from a sentence. The sentences used and the vocabulary are intentionally hand-designed to be simple. The attention block weights are hand-coded, instead of trained. The ANN FFN is trained. python import numpy as np np.set printoptions precision=2, suppress=True | Word | can fly? | has wheels? | can speak? | |---|---|---|---| | Rock | 0 | 0 | 0 | | Human | 0 | 0 | 1 | | Car | 0 | 1 | 0 | | Talking Tow Truck | 0 | 1 | 1 | | Crow | 1 | 0 | 0 | | Flying Superhero | 1 | 0 | 1 | | Plane | 1 | 1 | 0 | | Talking Planes | 1 | 1 | 1 | word to attributes = { "Rock": 0, 0, 0 , "Human": 0, 0, 1 , "Car": 0, 1, 0 , "Talking Tow Truck": 0, 1, 1 , "Crow": 1, 0, 0 , "Flying Superhero": 1, 0, 1 , "Plane": 1, 1, 0 , "Talking Planes": 1, 1, 1 , } for word, bits in word to attributes.items : print f"{word:20} {bits}" Rock 0, 0, 0 Human 0, 0, 1 Car 0, 1, 0 Talking Tow Truck 0, 1, 1 Crow 1, 0, 0 Flying Superhero 1, 0, 1 Plane 1, 1, 0 Talking Planes 1, 1, 1 The ANN's output - the logits probability scores for the next word is a correction to the input word. The input is fixed bit-encoding of 6 bits. Input word - 3 bits, one-hot, or IDs design choice Correction - 3 bits Output word - attributes after the correction attributes to word = {bits: word for word, bits in word to attributes.items } def apply correction word, correction : """Correction is 3 bits: 1 means flip that attribute.""" bits = word to attributes word out = tuple b ^ c for b, c in zip bits, correction return attributes to word out print apply correction "Rock", 1, 0, 0 add flight print apply correction "Human", 1, 0, 0 add flight Crow Flying Superhero Pseudocode - not run here. for loop in range num epochs : for batch inputs, batch targets in loader: prediction = model.forward batch inputs predict loss = CrossEntropyLoss prediction, target how far off? loss.backward gradient descent In the full workshop, the notebook with the full ANN implementation is available. Visit the relevant TechAarvam pages to locate them. Input is now a sentence: Crow keep-flight swap-speech he-is? Sentences have complex structure. Word order varies. So: fixed sentence structure, small defined vocabulary. Objects: Rock, Human, Crow, Flying superhero Actions: swap-flight, swap-speech, keep-flight, keep-speech Interrogative: he-is? 8-bit vectors, shown as 4 + 4. Bits 1-4 - attributes and correction: | Token | can fly | can speak | swap flight | swap speech | |---|---|---|---|---| | Rock | 0 | 0 | 0 | 0 | | Human | 0 | 1 | 0 | 0 | | Crow | 1 | 0 | 0 | 0 | | Flying superhero | 1 | 1 | 0 | 0 | | swap-flight | 0 | 0 | 1 | 0 | | swap-speech | 0 | 0 | 0 | 1 | | keep-flight | 0 | 0 | 0 | 0 | | keep-speech | 0 | 0 | 0 | 0 | | he-is? | 0 | 0 | 0 | 0 | Bits 5-8 - token type: | Token | object? | flight act? | speech act? | question? | |---|---|---|---|---| | Rock | 1 | 0 | 0 | 0 | | Human | 1 | 0 | 0 | 0 | | Crow | 1 | 0 | 0 | 0 | | Flying superhero | 1 | 0 | 0 | 0 | | swap-flight | 0 | 1 | 0 | 0 | | swap-speech | 0 | 0 | 1 | 0 | | keep-flight | 0 | 1 | 0 | 0 | | keep-speech | 0 | 0 | 1 | 0 | | he-is? | 0 | 0 | 0 | 1 | bit labels = "can fly", "can speak", "swap flight", "swap speech", "object?", "flight act?", "speech act?", "question?" token to bitvec = { fly spk swF swS obj flA spA q "Rock": 0, 0, 0, 0, 1, 0, 0, 0 , "Human": 0, 1, 0, 0, 1, 0, 0, 0 , "Crow": 1, 0, 0, 0, 1, 0, 0, 0 , "Flying superhero": 1, 1, 0, 0, 1, 0, 0, 0 , "swap-flight": 0, 0, 1, 0, 0, 1, 0, 0 , "swap-speech": 0, 0, 0, 1, 0, 0, 1, 0 , "keep-flight": 0, 0, 0, 0, 0, 1, 0, 0 , "keep-speech": 0, 0, 0, 0, 0, 0, 1, 0 , "he-is?": 0, 0, 0, 0, 0, 0, 0, 1 , } token to bitvec = {k: np.array v for k, v in token to bitvec.items } for tok, v in token to bitvec.items : print f"{tok:18} {v :4 } {v 4: }" Rock 0 0 0 0 1 0 0 0 Human 0 1 0 0 1 0 0 0 Crow 1 0 0 0 1 0 0 0 Flying superhero 1 1 0 0 1 0 0 0 swap-flight 0 0 1 0 0 1 0 0 swap-speech 0 0 0 1 0 0 1 0 keep-flight 0 0 0 0 0 1 0 0 keep-speech 0 0 0 0 0 0 1 0 he-is? 0 0 0 0 0 0 0 1 sentence = "Crow", "keep-flight", "swap-speech", "he-is?" X = np.stack token to bitvec t for t in sentence print X.shape X 4, 8 array 1, 0, 0, 0, 1, 0, 0, 0 , 0, 0, 0, 0, 0, 1, 0, 0 , 0, 0, 0, 1, 0, 0, 1, 0 , 0, 0, 0, 0, 0, 0, 0, 1 Skip the equations and go to the hand-written Q, K, V weights to get the idea behind these equations first. N heads. Each head does: python def softmax z, axis=-1 : z = z - z.max axis=axis, keepdims=True e = np.exp z return e / e.sum axis=axis, keepdims=True def head X, Wq, Wk, Wv : Q, K, V = X @ Wq, X @ Wk, X @ Wv d h = Wq.shape 1 scores = Q @ K.T / np.sqrt d h A = softmax scores return A @ V, A, Q, K, V Q - Query. K - Key. V - Value payload . The Q, K, V are intermediate tensors, what the model has are the corresponding weights. The weight matrices transform the input in three ways to get the Q, K, V . X is the input. Its the full sentence in our example. A context length full of tokens as input. The three weight matrices transform the input. The three matrices perform - three tasks. The weight matrices are per-head. i.e each head can extract out different information from the tokens. WQ is weights for getting the queries for this head. WK are the weights that transform the input tokens to keys. keys are like answers to the queries WV is the payload, if the question and the answer for a pair of tokens get a high score, the W V helps extract the payload or values from the tokens to pass along towards the FFN toward the concatenation operation, then the FFN . Every word asks a question; And also the question is not the same across heads. Each word in Each head can ask a different question; We have two heads in this hand-constructed example Object Head, Action Verb-Object head Object head - he-is? asks: are you an Object? Keys answer. Every token answers. Object head - "Crow" replies: I am an Object. he-is? in the object head asks - are you an object? Wq1 = np.zeros 8, 2 ; Wq1 7 = 8, 0 Wk1 = np.zeros 8, 2 ; Wk1 4 = 1, 0 Wv1 = np.zeros 8, 2 ; Wv1 0 = 1, 0 ; Wv1 1 = 0, 1 print "Wq1\n", Wq1, "\n\nWk1\n", Wk1, "\n\nWv1\n", Wv1 Wq1 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 8. 0. Wk1 0. 0. 0. 0. 0. 0. 0. 0. 1. 0. 0. 0. 0. 0. 0. 0. Wv1 1. 0. 0. 1. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. Wq2 = np.zeros 8, 2 ; Wq2 7 = 8, 8 Wk2 = np.zeros 8, 2 ; Wk2 5 = 1, 0 ; Wk2 6 = 0, 1 Wv2 = np.zeros 8, 2 ; Wv2 2 = 1, 0 ; Wv2 3 = 0, 1 print "Wq2\n", Wq2, "\n\nWk2\n", Wk2, "\n\nWv2\n", Wv2 Wq2 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 8. 8. Wk2 0. 0. 0. 0. 0. 0. 0. 0. 0. 0. 1. 0. 0. 1. 0. 0. Wv2 0. 0. 0. 0. 1. 0. 0. 1. 0. 0. 0. 0. 0. 0. 0. 0. Crow, bits 1-8: 1 0 0 0 1 0 0 0 The payload is Crow's object attributes. crow = token to bitvec "Crow" print "x Crow ", crow print "x Crow @ Wk1 ", crow @ Wk1, " <- key: I am an object" print "x Crow @ Wv1 ", crow @ Wv1, " <- value: can fly, can speak " x Crow 1 0 0 0 1 0 0 0 x Crow @ Wk1 1. 0. <- key: I am an object x Crow @ Wv1 1. 0. <- value: can fly, can speak swap-flight, bits 1-8: 0 0 1 0 0 1 0 0 The payload is the correction. sf = token to bitvec "swap-flight" print "x swap-flight ", sf print "x swap-flight @ Wk2 ", sf @ Wk2, " <- key: I am a flight action" print "x swap-flight @ Wv2 ", sf @ Wv2, " <- value: swap flight, swap speech " x swap-flight 0 0 1 0 0 1 0 0 x swap-flight @ Wk2 1. 0. <- key: I am a flight action x swap-flight @ Wv2 1. 0. <- value: swap flight, swap speech he-is? is the last token, so we read row 3 of the attention matrix. out1, A1, Q1, K1, V1 = head X, Wq1, Wk1, Wv1 out2, A2, Q2, K2, V2 = head X, Wq2, Wk2, Wv2 q = sentence.index "he-is?" print "Object head - attention from he-is?" for tok, k, a in zip sentence, K1, A1 q : print f" {tok:14} key={k} softmax={a:.2f}" print " head 1 output:", out1 q print "\nAction head - attention from he-is?" for tok, k, a in zip sentence, K2, A2 q : print f" {tok:14} key={k} softmax={a:.2f}" print " head 2 output:", out2 q Object head - attention from he-is? Crow key= 1. 0. softmax=0.99 keep-flight key= 0. 0. softmax=0.00 swap-speech key= 0. 0. softmax=0.00 he-is? key= 0. 0. softmax=0.00 head 1 output: 0.99 0. Action head - attention from he-is? Crow key= 0. 0. softmax=0.00 keep-flight key= 1. 0. softmax=0.50 swap-speech key= 0. 1. softmax=0.50 he-is? key= 0. 0. softmax=0.00 head 2 output: 0. 0.5 Two action words split the mass, so head 2 gives 0, 0.5 . WO scales by 2 - 0, 1 . Wo = np.diag 1.0, 1.0, 2.0, 2.0 head 2 mass was split across 2 words concat = np.concatenate out1 q , out2 q final = concat @ Wo print "concat ", concat print "after W O ", final print for name, val in zip bit labels :4 , final : print f" {name:12} {val:.0f}" concat 0.99 0. 0. 0.5 after W O 0.99 0. 0. 1. can fly 1 can speak 0 swap flight 0 swap speech 1 Hand-built weights pulled the object attributes and the correction attributes into 4 bits. Crow: flies, no speech. Correction: swap speech. - Flying superhero. Same 4 bits the ANN wanted. fly, speak, swap fly, swap speak = int v for v in final.round obj in = fly, 0, speak ANN bits: fly, wheels, speak correction = swap fly, 0, swap speak print "object in :", attributes to word obj in print "correction:", correction print "object out:", apply correction attributes to word obj in , correction object in : Crow correction: 0, 0, 1 object out: Flying Superhero This notebook is part of the support files for the TechAarvam workshop Build Your Own Model https://www.techaarvam.com/workshops/build-your-own-model . © TechAarvam. You are free to use, copy, modify, share and build on this material, including for commercial purposes, provided you credit TechAarvam and link back to