Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published 9 days ago • 144
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements Paper • 2608.17310 • Published 22 days ago • 108
yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF Text Generation • 12B • Updated Jun 19 • 521k • 2.88k
Running on Zero Agents Featured 2.22k Qwen3-TTS Demo 🎙 2.22k Generate speech from text with voice design, cloning, or presets
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices Paper • 2512.01374 • Published Dec 1, 2025 • 109