ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
Abstract
ARC improves fairness in group-based reinforcement learning for open-ended agents by conditioning rollout comparisons on strategy, enabling more context-appropriate behavior in responsive user-agent interaction.
Open-ended real-world interaction admits multiple valid behaviors: an agent may answer directly, ask for clarification, provide progress updates, or confirm before acting. This flexibility breaks a core assumption behind group-based RL: rollouts compared within a group are no longer guaranteed to be behaviorally comparable. As a result, reward-model preferences over interaction style can distort relative advantages and steer optimization toward reward-preferred behaviors rather than context-appropriate ones. We formalize this as a reward fairness problem and propose ARC (Advantage Regularization via Conditioning), a training recipe that restores fairer relative comparison through strategy-conditioned rollout grouping, together with hybrid rewards and entropy regularization. We study ARC in our proposed \inter, a novel paradigm for responsive, steerable, and execution-aware user-agent interaction that decouples user-visible communication from latent reasoning and tool use. \inter\ also provides the annotation and distillation pipeline for constructing \inter-86K, our strategy-annotated training corpus for supervised and RL training. Empirically, ARC substantially strengthens the core τ/τ^2 tool-use benchmarks, while \inter\ reduces time-to-first-token from 4.91s to 1.27s relative to a think-style baseline. Together, these results suggest that a central bottleneck in open-ended interactive learning is not only how agents are rewarded, but whether their behaviors are compared fairly in the first place. The ARC implementation and \inter-86K training data will be released.
Community
ARC: Fair Relative Advantage Comparison in Open-Ended Real-World Interaction
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- RISE-RL: Rubric-Informed Selective Exploration for Open-Ended Reinforcement Learning (2026)
- Off-Context GRPO: Learning to Reason on Hard Problems using Privileged Information (2026)
- Rubrics as Privileged Information for Open-Ended Generation (2026)
- Beyond Action Imitation: Learning a Decision-Aware User Simulator for Online Advertising (2026)
- RRPO: Reference-Relative Policy Optimization with Stratified Conditional Rollouts (2026)
- Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation (2026)
- Towards Better Agents for Multi-Turn User Interaction: The Next User Turn Is More Than Context (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.13622 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 1
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper