The user applied Opus 5.5 to optimize llama.cpp inference for the Swift Qwen 3.8 27B Q6_K model on an RTX 5090 GPU. This achieved a decode throughput of approximately 143 tokens per second and a prefill throughput of about 2840 tokens per second. Optimization details were illustrated with screenshots, though the exact recipe was not fully disclosed.

Read original