WhatAttention plus feed-forward is one block. GPT-2 small stacks 12 blocks, GPT-3 175B stacks 96.
HowEvery block reads the running list of numbers at each position and adds its result to it. Nothing is replaced, only added.
Why it mattersLater blocks build on what earlier blocks found. The last position's list after the top block is what gets scored.