FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation
Abstract
Long-horizon video generation requires models to effectively leverage an increasingly long generation history. As the generated history grows, retaining all previous content becomes increasingly expensive and redundant, making effective historical selection essential. Existing approaches often determine historical relevance based on the current content. However, information relevant to the present is not necessarily useful for future generation, while seemingly less relevant history may become important later. Our key insight is that historical information should be selected according to its relevance to future information needs. Capturing these needs does not require generating the full future; instead, a compact representation of what becomes important next is sufficient to guide historical selection. Building on this insight, we propose FrameMorrow, a prospective frame selector that predicts a small set of prospective tokens representing future information needs and uses them to identify relevant information from history. FrameMorrow selects explicit historical frames rather than model-specific internal states, enabling plug-and-play integration across diverse generators, including closed-source models, with little additional inference cost. We evaluate FrameMorrow across five benchmarks and 11 generative models spanning long-video generation, interactive generation, and action-conditioned world models. Extensive experiments demonstrate consistent improvements in long-range consistency, visual quality, and action alignment across diverse generation settings.
Community
We introduce FrameMorrow, a prospective frame selector for long-horizon video generation. Instead of selecting historical frames based only on their relevance to the present, FrameMorrow anticipates future information needs and uses them to identify useful history. It is plug-and-play across diverse video generators and consistently improves long-range consistency and visual quality across multiple benchmarks.
Good paper
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel →
Download all files from Hugging Face
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation (2026)
- LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation (2026)
- MosaiChunk: Compositing Spatio-Temporal Memory for Autoregressive Video Generation (2026)
- LongTake: Learning to Sustain Dynamics in Long-Horizon Video Generation (2026)
- WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory (2026)
- DensityKV: Density-Guided KV Cache Compression for Long Video Generation (2026)
- RECAP-Forcing: Retaining Content Appearances for Long Video Generation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.38839 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
