This paper presents a comparative study of softmax attention versus four recurrent linear-attention architectures: DeltaNet, Gated DeltaNet, Kimi Delta Attention, and Gated DeltaNet-2. By expressing these mechanisms through a common recurrent-memory notation, the authors analyze differences in expressivity, memory decay, and write/erase capabilities. The research aims to address the quadratic computational cost of standard self-attention in long-context training and inference.

Read original