A Reddit user asks the community which local large language model currently performs best for day‑to‑day coding and agent loops on consumer GPUs or Macs, comparing options like Qwen, Gemma, Llama, DeepSeek, and Mistral. The query highlights trade‑offs between tool‑calling reliability, raw token‑per‑second speed, and long‑context handling, and invites users to share any recent switches and their reasons.

Read original