Abstract
Large language models often solve independent problems in the same conversation. Can computation from earlier problems help them solve new ones? To answer this question, we first conduct preliminary experiments showing that retained history can raise or lower later-turn accuracy, even within the same domain. To understand these effects, we use controlled replay to isolate internal state changes specific to each problem-history pairing. Across different histories, these changes preserve similar relationships among current problems. To improve reasoning under retained history, we introduce STAIR (Stale-Token Attention for Inter-query Reuse). STAIR captures keys and values from earlier response generation in a fixed bank. It learns to redirect current queries when they read this bank during prompt processing. The base model remains frozen; only 12,288 parameters are trained. Across three Qwen models and four benchmarks, STAIR improves average later-turn accuracy by up to 11.67 percentage points over the unmodified model with history.
Community
Can computation from earlier problems help LLMs solve new ones? We show that retained history can both help and hurt reasoning after task switches. We introduce STAIR, which learns to re-address historical K/V states with only 12,288 trainable parameters while keeping the LLM frozen, improving later-turn accuracy by up to 11.67 percentage points.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- On-Demand Attention: Language Models Know When to Recall (2026)
- Settle: Learning When to Stop Reasoning (2026)
- RecurTrace: Adaptive Latent Reasoning with Loop-Time Memory (2026)
- OPSRD: On-Policy Self-Role Distillation (2026)
- AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models (2026)
- SANTA++: Sampling Attention through Representative Keys (2026)
- Delayed Supervision for Test-Time Language Models (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.39394 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper