The author implemented the Qwen3.5 architecture in FPGA fabric to enable INT4 quantized inference of 9B‑ and 27B‑parameter language models on inexpensive eBay‑sourced mining hardware. They selected the SQRL FK33 board, which provides 8 GB of HBM2 memory with approximately 400 GB/s bandwidth, priced around $280, and plan to scale using multiple cards.

Read original