VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Paper • 2607.27380 • Published 12 days ago • 71
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 14 days ago • 36
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs Paper • 2607.17423 • Published 22 days ago • 167
SWE-Pruner Pro: The Coder LLM Already Knows What to Prune Paper • 2607.18213 • Published 21 days ago • 79
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Paper • 2606.29538 • Published 25 days ago • 142
SearchEyes: Towards Frontier Multimodal Deep Search Intelligence via Search World Simulation Paper • 2607.05943 • Published Jul 7 • 15
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation Paper • 2603.08652 • Published Mar 9 • 40
OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every Iteration Paper • 2602.05400 • Published Feb 5 • 354