The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Paper • 2609.36484 • Published 13 days ago • 553
DeskForge: Dense Supervision from Desktop Environments for Computer-Use Agents Paper • 2610.02320 • Published 11 days ago • 19
RobotUse: Allocating Computation, Context, and Decisions Paper • 2610.04929 • Published 8 days ago • 19
LexReward: A Taxonomy-Driven Reward Framework for Legal Language Models Paper • 2609.39071 • Published 12 days ago • 66
Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL Paper • 2609.37200 • Published 13 days ago • 140
UniEvo-VL: An On-policy Self-Distillation Training Recipe for Multimodal Model Self-improvement Paper • 2609.38721 • Published 12 days ago • 274
InfoBayAI/CT-Scan-Radiology-Reports-Without-Findings-Dataset Viewer • Updated 3 days ago • 2.6k • 41 • 7
StanfordAIMI/stanford-deidentifier-with-radiology-reports-and-i2b2 Token Classification • Updated Nov 23, 2022 • 3.84k • 14
In-Context Learning for Robots: Methods and Applications Paper • 2609.36012 • Published 14 days ago • 322