EULLM is an open-source, Rust-based LLM inference engine under Apache 2.0 that is Ollama/OpenAI-API compatible and ships as a single binary with no Python or Docker dependencies. It runs on Windows, Linux, and macOS with CPU/CUDA/Metal support, and a recent test achieved 10 tok/s on a 35B MoE model using an ARM CPU without a GPU. Designed for EU-sovereign AI infrastructure, it features zero telemetry and a local-only audit trail for AI Act compliance.
Read original
reddit/r/LocalLLM