HippoCamp: Benchmarking Contextual Agents on Personal Computers Paper • 2604.01221 • Published Apr 1 • 30
FileGram: Grounding Agent Personalization in File-System Behavioral Traces Paper • 2604.04901 • Published Apr 6 • 40
Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published 21 days ago • 44
HumanCLAW: Can Vision-Language Models Act Through a Body? Paper • 2607.27180 • Published 9 days ago • 76
OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe Paper • 2511.16334 • Published Nov 20, 2025 • 96
HumanCLAW: Can Vision-Language Models Act Through a Body? Paper • 2607.27180 • Published 9 days ago • 76