The Qwen 3.8 27B model achieves deployment on a single RTX 5090 GPU using Ollama, utilizing approximately 22.4 GB of its 32 GB VRAM with a context length of around 175k tokens. While the setup performs well for daily coding tasks, it faces challenges sustaining long‑running agentic workflows compared to dedicated cloud solutions. This demonstration shows that a local 5090 can run this model efficiently for developers seeking self‑hosted alternatives.
Read original
reddit/r/LocalLLM