Papers
arxiv:2608.13552

PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives

Published on Aug 13
· Submitted by
Minghong Cai
on Aug 14
Authors:
,
,
,
,
,
,
,
,
,
,

Abstract

PlayWorld benchmarks interactive video world models by using multi-modal agents to pursue long-horizon objectives, evaluating geometry consistency, interaction fidelity, and state evolution.

Video world models simulate future states conditioned on current observations and user actions. Recent systems have demonstrated impressive video consistency and action controllability over long sequences. However, fairly comparing these interactive models remains challenging. In practice, a human player typically evaluates a world model by pursuing long-horizon objectives through interaction. For example, a user may turn around 360 degrees to see whether the environment remains consistent, or walk into the water and inspect whether realistic water ripples are generated. The action sequence required to achieve the same objective may vary substantially between models, making fixed action-conditioned evaluation unsuitable for cross-model comparison. To address this, we employ multi-modal Agent Players to interact with world models toward specified long-horizon objectives. Building on this paradigm, we introduce PlayWorld, a benchmark providing 171 scenarios, each with a specified objective. To evaluate performance thoroughly, we assess models along four core dimensions: geometry consistency, interaction fidelity, out-of-sight evolution, and insight evolution. In addition, we incorporate basic ability metrics for video quality and controllability. Experiments across nine state-of-the-art world models reveal that current models remain unreliable on long-horizon interactive objectives, particularly in maintaining spatial consistency and persistent state evolution. Code and data are available at https://github.com/kxding/PlayWorld.

Community

Paper submitter

PlayWorld, a new benchmark for interactive video world models. World models vary substantially in action granularity and response speed, so executing the same fixed action sequence often fails to bring them to the same state, making fair comparisons difficult. Instead of relying on predefined trajectories, PlayWorld introduces a multimodal Agent Player that interacts with each model in a closed loop—like a real player—while pursuing the same long-horizon goal. The agent continuously observes the generated frames and action history, dynamically adapting its actions throughout the interaction.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.13552
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.13552 in a model README.md to link it from this page.

Datasets citing this paper 1

Spaces citing this paper 1

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.