Abstract
We introduce the Extender, a log-structured variant of the standard Transformer architecture. In a standard Transformer, each layer communicates with subsequent layers exclusively via the residual h, a superposition channel. The Extender adds a concatenation channel x: each layer ell emits both a residual update δ_ell which is added to h, and a much smaller extension ε_ell which is appended to x. While both the FFN and q see h, the attention kv projections take only x as input. As a result, the fully extended x contains the complete input for the kv projections of all layers, reducing the persistent attention memory footprint from 2Ld_{model} to sum|ε_ell|. We find that with |ε_ell|=32, the Extender matches Transformer accuracy on short-context (CORE) tasks at 199M-924M parameters, and exceeds Transformer accuracy on long-context (RULER) workloads, again at 924M parameters. For our 1664-wide, 924M model, the Extender's persistent attention memory footprint is 104times smaller than MHA. The memory savings grow with model width.
Community
This variation on the Transformer reduces the attention memory footprint between turns to 32 features per layer, while retaining (CORE) or improving (RULER) accuracy for our tested models with 200/450/900M parameters. This is before sparse attention, quantization, or layer reuse.
If confirmed on larger model sizes, the Extender could result in memory savings in excess of 50x.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Routed Graph Handoff: Adaptive Format Selection for Multi-Agent LLM Delegation (2026)
- Certificates for short extending words in a finite automaton (2026)
- Nearly Optimal Strong Coresets for $\ell_p$ Subspace Approximation (2026)
- Improving TensorSketch Using Complex Random Variables (2026)
- Twisted Bracelets for Sorting by Transpositions: the Transposition Diameter of $S_{16}$ (2026)
- Certifiably Interpretable Training of ReLU-MLPs for Boolean Tasks with Guaranteed Truth-Table Generalization (2026)
- Improved Quantum Random Self-Reduction for Linear Problems (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.32759 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper