I ran a 110B LLM on 16GB of RAM. Here's the equation that predicts any model's speed on your machine
Article automatically generated from technical news.
My 2016 desktop — 16 GB RAM, SATA SSD — ran GLM-4.5-Air, a 110B-parameter model, streamed from disk. One equation predicted the speed before I pressed enter: 0.2-0.3 tok/s. It measured 0.19. That equation (tok/s = eta(tier) x bandwidth / active-bytes-per-token) is what this post hands you, plus the 30-minute probe that finds where your model breaks. But first, the control experiment that proves placement is the whole gam
Fonte originale