A user is seeking advice on optimizing local LLM performance for coding tasks on a MacBook Pro equipped with an M4 Max chip and 36GB of RAM. They are currently running a 4-bit quantized Qwen 3.5 27B model via LM Studio, achieving generation speeds of 10-15 tokens per second. The user aims to determine if this model is the most capable option for managing a large Python and React/TypeScript codebase given their hardware constraints.
Read original
reddit/r/LocalLLM