Task-CoEvolve: Efficient Harness Optimization via Adaptive Validation Task Selection Paper • 2608.20169 • Published 4 days ago • 10
360CityArena: A Realistic Virtual Urban Navigation Benchmark for Embodied Agents Paper • 2608.08814 • Published 19 days ago • 10
Paper Reconstruction Evaluation: Evaluating Presentation and Hallucination in AI-written Papers Paper • 2604.01128 • Published Apr 1 • 15
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction Paper • 2512.14620 • Published Dec 16, 2025 • 2
JMMMU-Pro: Image-based Japanese Multi-discipline Multimodal Understanding Benchmark via Vibe Benchmark Construction Paper • 2512.14620 • Published Dec 16, 2025 • 2
Jr. AI Scientist and Its Risk Report: Autonomous Scientific Exploration from a Baseline Paper Paper • 2511.04583 • Published Nov 6, 2025 • 5
ShinkaEvolve: Towards Open-Ended And Sample-Efficient Program Evolution Paper • 2509.19349 • Published Sep 17, 2025 • 2