A developer has reverse-engineered the Apple Neural Engine kernel to create an inference engine capable of running Qwen 3.8 27b FP16 with a 256k context window. The implementation achieves 7-8 tokens per second while consuming only 7 watts of power, keeping the GPU idle. This optimization allows for high-tier model inference on portable devices with significantly reduced battery drain.

Read original