The paper introduces Persistent State Machines (PSM) that employ INT4‑compressed in‑memory cells to implement attention mechanisms in large language models, enabling persistent state across layers. This design reduces memory bandwidth and improves inference efficiency while maintaining model performance. Empirical results show notable speedups and lower memory usage compared to traditional attention approaches.
Read original
hackernews