The post notes emergence of highly specialized inference engines such as Strata, ninfer, DwarfStar, Splash, llamAmpere, and gufo, which sacrifice the broad model support of llama.cpp and vLLM to optimize for a limited set of models and specific hardware like Strix Halo. This trend suggests a bifurcation where general‑purpose runtimes maintain compatibility while disposable, over‑fit engines deliver peak performance for targeted workloads. The author predicts this split will become the standard approach moving forward.

Read original