Concepts explained
Where Memory Lives — DRAM and HBM
Today's computers do not jam on calculation. They jam on fetching memory. Why computer memory is a bucket that never stops leaking, and why HBM — those buckets stacked upward beside the processor — became the key part of the AI age.
Memory comes in two kinds
A computer has two kinds of remembering part. One forgets when the power goes but is very fast — a notebook open on the desk. That is memory, RAM.
The other keeps its contents with the power off but is slow — books on the shelf. That is storage, the SSD. Photos and apps live there and are brought to the desk only when used.
Memory is a leaking bucket
Look inside the desk memory and it is astonishingly simple. Each cell is one tiny bucket and one lid. Charge in the bucket means 1; empty means 0. The lid is the switch from piece ten.
The jam is not in the calculating
Here is this piece's surprise. Piece five showed the processor's sense of time: if one beat is a second, a trip to memory takes minutes.
Piece six called the GPU thousands of helpers. But however many helpers, no potatoes can be peeled if none arrive from the store. Piece seven's AI must read billions of knobs from memory every time it calculates. The multiplying itself is instant. Fetching eats all the time.
Stack the warehouse, put it beside the factory
There were two fixes: bring the warehouse closer, and widen the road between warehouse and factory.
Memory chips laid side by side on the board, joined to the processor by long wires. A distant warehouse and a few trucks.
Memory chips stacked a dozen high, the floors joined by thousands of tiny vertical holes, the whole tower set right beside the processor. A twelve-storey warehouse next to the factory with a thousand conveyor belts.
HBM means 'memory with a wide road'. It carries dozens of times more at once than before. That is why AI GPUs have several HBM towers pressed right up against them — the device that keeps piece six's thousands of helpers supplied with potatoes.
A little further in
Why not just anyone can make it
Stacking is easy to say and hard to do. Each memory chip must be ground thinner than paper; thousands of holes finer than a hair must be drilled through it to join top to bottom; a dozen must be stacked without the slightest misalignment. One skewed layer ruins the tower.
And there is heat. Piece eight's data centres fight heat, remember. Stack chips and the middle floors have nowhere to send theirs.
Closing the pair
Pieces ten and eleven fill the gap left in piece one. The switch is made from sand, printed with light; give it a lid and a leaking bucket holds memory. When carrying, not calculating, became the bottleneck, the buckets were stacked upward.
What remains is the story that breaks the switch's own rule — the switch that is neither on nor off. On to the quantum computer.
The question that remainsIf fetching memory is slower than calculating, what is it that we call 'the speed of a computer' the speed of?