Strata enables the 125‑billion‑parameter Qwen3.8‑Flash‑Next model to run on consumer PCs by distributing its 24,576‑expert mixture across GPU, RAM and SSD. On an NVIDIA RTX 5070 (12 GB VRAM) the Q2_0 quantization generates text at ~94 tokens/s and ingests prompts at ~2,650 tokens/s, while the IQ3_S variant yields ~53 tokens/s generation and ~1,620 tokens/s prompt processing. On an AMD RX 9070 XT (16 GB VRAM) the corresponding speeds are ~60 tokens/s generation and ~1,160 tokens/s input for Q2_0. A RTX 3090 (24 GB VRAM) is projected to achieve 100–140 tokens/s generation. Memory requirements scale with quantization: 32 GB RAM suffices for the Coder variant, 48 GB for IQ2_XS/Q2_0, and 64 GB or more allows all sizes including IQ3_S. The installer downloads ~70 GB of model data, allocates 35–55 GB of RAM (part locked for the GPU), and uses an SSD for a large lookup table, enabling long‑context handling (up to 128K tokens) at >1,000 tokens/s chunk reads. Strata supports multi‑GPU sharing, Windows 10/11 or Linux, and provides an OpenAI‑compatible API endpoint at http://127.0.0.1:8080/v1 for integration with coding agents.
Read original
hackernews