While NVIDIA's H100 provides 15× more FP16 compute TFLOPS than the T4, a 7B parameter LLM exhibits a much larger performance gap, jumping from ~15 tok/s to over 3,500 tok/s. This indicates that raw TFLOPS alone do not fully explain the massive increase in inference throughput seen on the H100, especially when utilizing continuous batching.

Read original