A user evaluated the coding performance of Qwen 3.6-27B (Q4_K_M), Ornith-9B-Q8, and a custom fine-tune on an RTX 4060 Ti 16GB system. Using Ollama and llama.cpp, the Qwen 3.6-27B model achieved a 63.6% resolution rate (14/22) on SWE-bench benchmarks. The author is seeking recommendations for more efficient coding models compatible with their current hardware constraints.
Read original
reddit/r/LocalLLM