RLHND: Video Foundation Models as Physically Grounded Hand Trackers for Robot Learning Paper • 2610.09455 • Published 4 days ago • 38
AdvSim2Real : Training Web Agents Against Adaptive Prompt Injection in a Web World Model Paper • 2610.08773 • Published 4 days ago • 15
Base Models Can Reason By Taking a Cue From Training Data Paper • 2610.06851 • Published 6 days ago • 24
HLA-WM: Hybrid Linear Attention for Long-Horizon Video World Models Paper • 2610.05739 • Published 6 days ago • 21
World Editing: Intervening on Executable Worlds at Increasing Depth Paper • 2610.02331 • Published 10 days ago • 30
In-Distribution Forcing for Long Video Generation at Test Time Paper • 2610.03120 • Published 9 days ago • 49
ProgressCompass: Embodied Progress Reward Models Are Lost Without the Right Context Paper • 2609.36684 • Published 12 days ago • 20
Ego2Act: Evaluating Goal-Directed Manipulation in Egocentric Video Generation Paper • 2610.01092 • Published 10 days ago • 34
AutoGUIWorld: Image Generators as Visual World Models for GUI Agent Paper • 2610.01215 • Published 10 days ago • 63
Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL Paper • 2609.37200 • Published 12 days ago • 140
Beyond Memory: Harnessing Long-Horizon Agents with Explicit Belief States Paper • 2610.01415 • Published 10 days ago • 91
World Observer: Joint Actor-Observer Generation for Persistent World Modeling Paper • 2610.02162 • Published 10 days ago • 89