The user repurposed an iPhone 17 Pro Max as a secondary GPU for a 24 GB M4 Pro MacBook, offloading part of the Qwen 3.8 27B model to the phone. This setup yields 29‑44% faster prefill performance while the iPhone stores a segment of the model’s context window. The cited speedup reflects end‑to‑end prefill rate after correcting the phone‑side measurement.
Read original
reddit/r/LocalLLaMA