The benchmark found that llama.cpp, llamafile, LM Studio, and Ollama deliver nearly identical prompt‑processing throughput on a Mac Studio (M4 Max Metal), a Steam Deck (Vulkan), and a Linux box with an NVIDIA L40S (CUDA), staying within a few percent of each other when the model and environment are fixed. The observed tens‑of‑percent performance variations stem from how each server is built, not from differences in the underlying inference engine.

Read original