MIMESIS: Learning User Simulators as Training Environments for Interactive Agents Paper • 2610.09484 • Published 5 days ago • 23
InvisMark: Invisible and Robust Watermarking for AI-generated Image Provenance Paper • 2411.07795 • Published Nov 19, 2024
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents Paper • 2602.06855 • Published Feb 6 • 83
AIRS-Bench: a Suite of Tasks for Frontier AI Research Science Agents Paper • 2602.06855 • Published Feb 6 • 83
VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice Paper • 2601.05175 • Published Jan 8 • 38
VideoAuto-R1: Video Auto Reasoning via Thinking Once, Answering Twice Paper • 2601.05175 • Published Jan 8 • 38
What Does It Take to Be a Good AI Research Agent? Studying the Role of Ideation Diversity Paper • 2511.15593 • Published Nov 19, 2025 • 59
Souper-Model: How Simple Arithmetic Unlocks State-of-the-Art LLM Performance Paper • 2511.13254 • Published Nov 17, 2025 • 140
Chain of Natural Language Inference for Reducing Large Language Model Ungrounded Hallucinations Paper • 2310.03951 • Published Oct 6, 2023
FactCG: Enhancing Fact Checkers with Graph-Based Multi-Hop Data Paper • 2501.17144 • Published Jan 28, 2025 • 5
Beyond Reasoning Gains: Mitigating General Capabilities Forgetting in Large Reasoning Models Paper • 2510.21978 • Published Oct 24, 2025 • 17
ATLAS: Decoupling Skeletal and Shape Parameters for Expressive Parametric Human Modeling Paper • 2508.15767 • Published Aug 21, 2025 • 18
Beyond Correlation: Interpretable Evaluation of Machine Translation Metrics Paper • 2410.05183 • Published Oct 7, 2024 • 1