The author built a home LLM inference server using dual 20GB RTX 3080 GPUs, totaling 40GB VRAM, to run Qwen 27B. The system includes an Intel i9-10900X CPU, 64GB DDR4 RAM, and an X299 motherboard with multiple PCIe slots. The total build cost was kept under $2000, with the GPUs sourced from Alibaba after filtering out scam sellers.

Read original