The author ran llama.cpp inference on a Galaxy S24 Ultra (Snapdragon 8 Gen 3) using non‑root Termux, first getting the Adreno 750 OpenCL backend functional, though it was slower than CPU execution. They then accessed the Hexagon v75 HTP through Qualcomm’s vendor/FastRPC interface, enabling source‑built HMX/HVX kernels for LLM matrix operations.

Read original