Running on Zero MCP 5 Valen Visual Decisions 👁 5 Visual question answering with candidate probabilities
Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games Paper • 2606.19338 • Published Jun 17 • 52
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation Paper • 2605.10912 • Published May 11 • 37