JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published 6 days ago • 89
Reinforcing Dual-Path Reasoning in Spatial Vision Language Models Paper • 2606.17539 • Published Jun 16 • 15
OmniGameArena: A Unified UE5 Benchmark for VLM Game Agents with Improvement Dynamics Paper • 2606.09826 • Published Jun 8 • 19
Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing Paper • 2606.05172 • Published Apr 16 • 1
LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation Paper • 2606.02553 • Published Jun 1 • 20
LongLive-2.0: An NVFP4 Parallel Infrastructure for Long Video Generation Paper • 2605.18739 • Published May 18 • 116
Nemotron 3 Nano Omni: Efficient and Open Multimodal Intelligence Paper • 2604.24954 • Published Apr 27 • 26
Agentic World Modeling: Foundations, Capabilities, Laws, and Beyond Paper • 2604.22748 • Published Apr 24 • 232
Perflow-Shuai/Wan2.1-T2V-1.3B-NonAR-DMD-4Step-LoRA-r64-iter1600-Teacher14B Text-to-Video • Updated 11 days ago
Perflow-Shuai/Wan2.1-T2V-1.3B-NonAR-DMD-4Step-LoRA-r64-iter1600-Teacher14B Text-to-Video • Updated 11 days ago
Perflow-Shuai/Wan2.2-5B-CFG5-to-CFG1-50Step-LoRA-r64-iter250 Image-to-Video • Updated 11 days ago • 25 • 1
Perflow-Shuai/Wan2.2-5B-CFG5-to-CFG1-50Step-LoRA-r64-iter250 Image-to-Video • Updated 11 days ago • 25 • 1
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 14 days ago • 36