How does a language model remember what was said? A transformer keeps a key and a value for every token it has read, in every attention layer, so its memory grows with every word. State-space and linear-attention layers do it the other way: Mamba-2, from Tri Dao and Albert Gu, and Gated DeltaNet, from Songlin Yang, Jan Kautz and Ali Hatamizadeh, fold everything they have read into a state of fixed size and update it in place. The models people actually run in 2026 mix the two. NVIDIA's Nemotron-3-Nano-4B has 21 Mamba-2 mixers and 4 attention layers. Qwen3.5-9B is Gated DeltaNet in three layers of every four. Both architectures run on Ferric, the institute's runtime, checked against their authors' own code, Nemotron-3-Nano-4B itself and Qwen3.5 at its 0.8B size, and both configs are public, so the memory ledger is arithmetic. Build it and something surprising appears. The recurrent layers are the cheap part, 85.03 MB and 52.69 MB however long the conversation runs. The few attention layers overtake them within a few thousand tokens and carry nearly all of a long one. That is why the 2026 frontier competes on what a fixed state does rather than on how big it is. Mamba-3 points out that decoding one token does about 2.5 operations for every byte it reads, against about 295 an H100 can do, so the work is waiting on memory, and its multi-input variant does more useful math per byte without growing the state. Gated DeltaNet-2 holds the state at exactly 262,144 numbers per layer, the same as the models it beats, and changes only how the state is erased and written. When decoding waits on memory, the bytes it reads set most of what it costs, so a fixed state that is used better is the energy story.