Multilingual GSM-Symbolic: What determines capability transfer across languages? Paper • 2610.03367 • Published 5 days ago • 47
Spatial Memory Intelligence: Endowing World Models with Understanding-Driven Long-Term Memory Paper • 2610.02521 • Published 6 days ago • 49
EditHero: A Benchmark for Long-Horizon Part-Level 3D Editing and Vibe Modeling Paper • 2610.02298 • Published 6 days ago • 50
Latent-MOPD: Latent Multi-Teacher On-Policy Distillation Paper • 2610.02381 • Published 6 days ago • 63
On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training Paper • 2609.36659 • Published 8 days ago • 81
Native Action-Prior Learning from Videos for World Action Models Paper • 2610.03391 • Published 5 days ago • 81
World Action Modeling with Progressive Visual Planning Paper • 2610.02508 • Published 6 days ago • 83
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite Paper • 2610.02826 • Published 5 days ago • 91
RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 6 days ago • 268
PixelUMM: Encoder-Free Unified Image and Video Understanding and Generation Paper • 2609.38597 • Published 8 days ago • 32
Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation Paper • 2610.01092 • Published 6 days ago • 33
DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence Paper • 2609.39222 • Published 7 days ago • 42
It's Not What the Image Shows: Irrelevant Context Destabilises VLM Judges Without Informing Them Paper • 2609.37863 • Published 8 days ago • 38
AutoDataBench: A Data-centric Testbed for Accelerating Auto Research Paper • 2609.40097 • Published 7 days ago • 45
Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models Paper • 2609.33355 • Published 10 days ago • 50
4Director: Controlling Video World Models with Rigid 3D Geometry Paper • 2610.02160 • Published 6 days ago • 38
SILSA: Sliding-Window Slice Latents for Topology-Preserving High-Resolution 3D Generation Paper • 2610.02201 • Published 6 days ago • 36
FrameMorrow: Future-guided Frame Selection with Prospective Tokens for Long-Horizon Video Generation Paper • 2609.38839 • Published 7 days ago • 93
Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training Paper • 2609.40111 • Published 7 days ago • 50