Papers
arxiv:2609.38169

STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization

Published on Sep 29
ยท Submitted by
Haokun Lin
on Oct 8
#1 Paper of the day
Authors:
,
,
,
,
,
,
,

Abstract

Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQuant allocates precision according to error magnitude and memory lifetime, and jointly fits key-row and value-column scales based on state distributions and key-row impact on output error. Experiments on Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct across both long- and short-generation benchmarks show that STEPQuant closely matches FP32-state accuracy under a nominal 6-bit budget and outperforms uniform INT8 in its 4-bit configuration. Integrated into SGLang with optimized GPU kernels, 6-bit STEPQuant achieves over 5x recurrent-state compression and reduces total serving memory by up to 68.7%. Our code is available at https://github.com/Dreamer-Toby/STEPQuant.

Community

๐Ÿš€ Excited to share STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State Quantization!

STEPQuant enables efficient low-bit quantization of recurrent states in linear attention by addressing quantization errors across both temporal and spatial dimensions.

โœจ Key highlights:

  • Near-FP32 accuracy with 6-bit states
  • Over 5ร— recurrent-state compression
  • Up to 68.7% reduction in total serving memory
  • Efficient SGLang implementation

๐Ÿ“„ Paper: https://arxiv.org/abs/2609.38169

๐Ÿ’ป Code: https://github.com/Dreamer-Toby/STEPQuant

Feedback and discussions are very welcome!

fig1

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.38169
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.38169 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.38169 in a dataset README.md to link it from this page.

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.