RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations Paper • 2610.01780 • Published 11 days ago • 282
4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes Paper • 2610.03715 • Published 10 days ago • 33
LVMT: Video Mask Transformer for Long-term Video Segmentation Paper • 2609.34895 • Published 13 days ago • 19
Scaling Properties of Same-Family On-Policy Distillation Paper • 2609.32722 • Published 16 days ago • 326
PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers Paper • 2609.32429 • Published 16 days ago • 7
AutoDataBench: Can Agents Write the Data That Feeds the Self-Improvement Loop? Paper • 2609.35025 • Published 14 days ago • 9
WISE-ATTA: When to Ask for Labels in Budgeted Active Test-Time Adaptation Paper • 2609.37687 • Published 13 days ago • 7