S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published 5 days ago • 32
Understanding Evolution Strategies for LLM Reasoning: Broader Reasoning Coverage than GRPO Paper • 2608.27351 • Published 9 days ago • 22
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts Paper • 2608.20061 • Published 16 days ago • 46
Zetta ζ: An Efficient Closed-Loop Embodied Harness for Self-Evolving Physical Intelligence Paper • 2608.16590 • Published 19 days ago • 150
EnvHarness: Awakening Static Worlds for Agent Learning Paper • 2608.19880 • Published 16 days ago • 274
Co-RL: Unsupervised Reasoning Emerges from Diverse Cohort in Multi-agent RL Paper • 2608.17253 • Published 17 days ago • 96
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements Paper • 2608.17310 • Published 18 days ago • 108
LuffyTheFox/Qwen3.6-35B-A3B-Uncensored-Genesis-Hermes-V13-GGUF Image-Text-to-Text • 35B • Updated about 8 hours ago • 929k • 594
view article Article What We Learned by Reproducing 2,200 papers from ICML abidlabs • 23 days ago • 110