Mirai S from Mirai Labs delivers up to 70 tokens per second with a 262k token context window on GPUs with at least 12 GB VRAM, incorporating an agentic workflow patch for 27‑parameter quant models. Users with 12 GB+ GPUs achieve full context and speed, while 16 GB+ GPUs gain additional VRAM placement for further performance gains.
Read original
reddit/r/LocalLLM