PSSA is a 1.5‑million‑parameter language model implemented from scratch in Rust without any deep‑learning framework. Its core is a selective diagonal state‑space layer (d_m=256 channels, d_s=16 states per channel) that propagates a fixed‑size recurrent state left‑to‑right, giving O(L) per‑token cost instead of the O(L²) quadratic cost of self‑attention. The layer feeds a hyperbolic episodic memory bank of 512 slots (Poincaré ball, key width 32) whose top‑4 nearest entries are retrieved with a softmax over distance, gated by a learned SiLU adapter. Plastic weights are updated online: novel states trigger write‑in slots with a refractory counter, and fast adapter weights are periodically consolidated into the base transition matrix via closed‑form ridge regression. Training on cleaned WikiText‑103 (12.7 M tokens) with identical optimizer schedules yields a cross‑entropy of 3.98 nats (perplexity 53.7) for PSSA versus 4.43 nats (perplexity 83.7) for a matched transformer, a 0.45 nat gap that persists on held‑out data (3.997 vs 4.429 nats, 24.1% vs 18.0% next‑token accuracy). Generation of 200 tokens on the same CPU runs in 226 ms for PSSA versus 2 735 ms for the transformer (~12× speedup). Gradient checks against a scalar reference path agree to ≤3×10⁻⁸.
Read original
hackernews