A new Rust-based inference framework has been developed that runs Qwen3.5 2B with vision-language (VL) support up to 10x faster than PyTorch on Apple Silicon. The framework natively supports text-to-speech (TTS), automatic speech recognition (ASR), optical character recognition (OCR), and GGUF model formats out of the box. Built for efficient deployment on Apple hardware, it aims to provide high-performance inference across multiple modalities without relying on Python-based ecosystems. Read original
reddit/r/machinelearningnews