Perplexity has open-sourced Lily, a specialized local inference engine developed in Rust with hand-written Metal kernels. Unlike general-purpose runtimes, Lily is optimized specifically for the Qwen3.6-35B-A3B model on Apple silicon, bypassing PyTorch and MLX to eliminate generality bottlenecks.

Read original