From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health Paper • 2609.25186 • Published 6 days ago • 26
1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation Paper • 2609.24432 • Published 6 days ago • 16
RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning Paper • 2609.20784 • Published 10 days ago • 56
ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models Paper • 2609.18487 • Published 11 days ago • 47
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 12 days ago • 73
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 13 days ago • 214
PhysStream: Streaming Physics-Grounded Video Generation with Structured Scene Memory and Fine-Grained Motion Control Paper • 2609.17521 • Published 12 days ago • 10
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training Paper • 2609.15051 • Published 13 days ago • 14
EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents Paper • 2609.05903 • Published 22 days ago • 65
NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction Paper • 2609.10715 • Published 18 days ago • 329
Scores Alone Do Not Prove Discovery: The Discovery Certification Protocol for Auditing AI Research Agents Paper • 2609.09219 • Published 20 days ago • 16
RoboTok: An Internet-Scale Data Engine for Human Demonstration Retrieval and Dexterous Manipulation Learning Paper • 2609.03199 • Published 25 days ago • 31
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published 24 days ago • 84
SynthGait-19K: A Physically Grounded Synthetic Video Dataset for Gait Parameter Estimation Paper • 2609.08108 • Published 19 days ago • 18
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience Paper • 2609.03241 • Published 24 days ago • 54
τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction Paper • 2609.04611 • Published 23 days ago • 11
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 24 days ago • 245
Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling Paper • 2608.30821 • Published 27 days ago • 67
Agent-G^2: Gaussian Guidance for Agentic Reinforcement Learning Paper • 2608.23318 • Published Aug 24 • 32