
AI Inference Chips: Why Memory Bandwidth Beats Peak FLOPS
Decode reads the whole model to produce one token. Working from NVIDIA's own published rack figures, here is the arithmetic that explains why serving throughput ignores the number on the box.
Tag

Decode reads the whole model to produce one token. Working from NVIDIA's own published rack figures, here is the arithmetic that explains why serving throughput ignores the number on the box.