An open‑source inference engine automatically compiles and tunes its kernels on the user's hardware, achieving up to 2× speedup over llama.cpp for running open models. It supports Apple Silicon, NVIDIA and AMD GPUs as well as CPU‑only environments. The project was highlighted in a Reddit post on r/LocalLLaMA.

Read original