MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement Paper • 2610.11959 • Published 4 days ago • 72
TokenRouter: Efficient Serving System for Token-Level LLM Routing Paper • 2610.12242 • Published 4 days ago • 130
WorldPlay2: Extending Real-Time Interactive World Models in Control and Horizon Paper • 2609.35560 • Published 14 days ago • 31
Learning to Learn from Context: Synthetic Training from Perturbed Public Documents Paper • 2609.33642 • Published 15 days ago • 33
Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning Paper • 2609.33781 • Published 15 days ago • 46
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning Paper • 2609.35767 • Published 14 days ago • 50
Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge Paper • 2609.34327 • Published 14 days ago • 41
CompoWorld: Compositional Environment Scaling for General Agents Paper • 2609.33665 • Published 15 days ago • 41
EmbodiedMemory-Bench: Benchmarking Embodied Memory for Long-Horizon Embodied Tasks Paper • 2609.28236 • Published 19 days ago • 39
Improving Test-Time Scaling with Adaptive Looped Transformers Paper • 2609.35748 • Published 14 days ago • 61
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published Sep 10 • 174
SecOPD: Mitigating Adaptive Prompt Injections by On-Policy Distillation Paper • 2608.21500 • Published Aug 21 • 41
AstraFlow: Dataflow-Oriented Reinforcement Learning for Agentic LLMs Paper • 2605.15565 • Published May 15 • 16
The Latent Space: Foundation, Evolution, Mechanism, Ability, and Outlook Paper • 2604.02029 • Published Apr 2 • 111
OpenResearcher: A Fully Open Pipeline for Long-Horizon Deep Research Trajectory Synthesis Paper • 2603.20278 • Published Mar 17 • 103
SkillsBench: Benchmarking How Well Agent Skills Work Across Diverse Tasks Paper • 2602.12670 • Published Feb 13 • 67
RLinf-USER: A Unified and Extensible System for Real-World Online Policy Learning in Embodied AI Paper • 2602.07837 • Published Feb 8 • 57