Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper ⢠2607.08964 ⢠Published 14 days ago ⢠74
OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification Paper ⢠2606.01476 ⢠Published May 31 ⢠9
Conditional Hypothesis Generation for LLM-Based Text Analysis with Researcher-Specified Covariates Paper ⢠2606.03029 ⢠Published Jun 2 ⢠6
Synthetic Sandbox for Training Machine Learning Engineering Agents Paper ⢠2604.04872 ⢠Published Apr 6 ⢠14
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Paper ⢠2509.00676 ⢠Published Aug 31, 2025 ⢠85
Self-Rewarding Vision-Language Model via Reasoning Decomposition Paper ⢠2508.19652 ⢠Published Aug 27, 2025 ⢠85
Skill Discovery for Software Scripting Automation via Offline Simulations with LLMs Paper ⢠2504.20406 ⢠Published Apr 29, 2025 ⢠8
MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning Paper ⢠2506.05523 ⢠Published Jun 5, 2025 ⢠34
SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement Paper ⢠2504.07934 ⢠Published Apr 10, 2025 ⢠21