The K2 Horizon 36B‑A4B language model contains 36 billion parameters in total but activates only 4 billion parameters per token during inference. This sparse activation dramatically lowers the computational and memory footprint, making the model suitable for local AI deployment on consumer‑grade hardware. The approach retains the expressive power of a large model while improving efficiency.
Read original
medium