Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation in Transformers Paper • 2608.15062 • Published 2 days ago • 3
Hybrid Quantum-inspired Kolmogorov-Arnold Networks for Privacy-Aware Federated Biosignal Learning Paper • 2608.13914 • Published 13 days ago • 3
Let's Scale Step by Step: Compute-Efficient Hyperparameter Transfer for Large-Scale Mixture-of-Experts Paper • 2608.20061 • Published 7 days ago • 43
FlashPrefill V2: Block-Sparse Prefill Attention for Long-Context LLM Serving Paper • 2608.19758 • Published 7 days ago • 18
AVA-Encoder: Towards Agent-Native Video Representation Learning Paper • 2608.12313 • Published 15 days ago • 42
ASI-Bench: At the Dawn of Artificial Superintelligence Paper • 2608.17271 • Published 9 days ago • 62
Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published 10 days ago • 149
MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling Paper • 2608.14783 • Published 13 days ago • 19
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Paper • 2608.12149 • Published 15 days ago • 30
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 17 days ago • 342
Skaling: Chinchilla's Exponents Meet Kaplan's Coupling Paper • 2608.07222 • Published 20 days ago • 10
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Paper • 2607.24027 • Published Jul 27 • 37