The author benchmarked four open decision models—Laya, Liquid's D1, Cloudflare's Clef‑Flash, and Interfaze's LEV—on an RTX 4090 by having each process nine Wikipedia articles about centipedes (≈9.5 k words) and flag centipede‑naming tokens, making one API call per word. The test aimed to measure relative inference speed despite differing model sizes. Results were shared to show how the models compare on identical hardware.
Read original
reddit/r/LocalLLaMA