FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis Paper • 2608.18580 • Published 8 days ago • 120
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models Paper • 2608.16887 • Published 10 days ago • 33
Intern-S2-Preview: Scientific Agentic Foundation Model Paper • 2608.13505 • Published 14 days ago • 70
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published 21 days ago • 62
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging Paper • 2608.03316 • Published 23 days ago • 26
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published 23 days ago • 103
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Paper • 2608.02602 • Published 24 days ago • 81
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Paper • 2607.13705 • Published Jul 15 • 45
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning Paper • 2607.07708 • Published Jul 8 • 88
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published Jul 8 • 64
Scaling the Horizon, Not the Parameters: Reaching Trillion-Parameter Performance with a 35B Agent Paper • 2606.30616 • Published Jun 29 • 105
Beyond the Current Observation: Evaluating Multimodal Large Language Models in Controllable Non-Markov Games Paper • 2606.19338 • Published Jun 17 • 51
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion Paper • 2605.23902 • Published May 22 • 47
ACC: Compiling Agent Trajectories for Long-Context Training Paper • 2605.21850 • Published May 21 • 61
OcclusionFormer: Arranging Z-Order for Layout-Grounded Image Generation Paper • 2605.21343 • Published May 20 • 8