Learn What's Left, Not What's Mastered: Saturation Aware Advantage Reweighting for Multi-Reward Policy Optimization Paper • 2608.16072 • Published 5 days ago • 147
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published 12 days ago • 338
TARS: Timestep-Aware Data Scaling for 3D-Free Video Re-Shooting Paper • 2607.28261 • Published 23 days ago • 116
stefanocarrera/sqlautophagycode_D_test_Qwen3-8B_t1.25_g5_run0_metrics Viewer • Updated Jul 20 • 579 • 37 • 1
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget Paper • 2607.14952 • Published Jul 16 • 212
EgoSteer: A Full-Stack System Towards Steerable Dexterous Manipulation from Egocentric Videos Paper • 2607.09701 • Published Jun 21 • 17
timaeus/rl-lm-pythia1b-sentiment-neg-alpha0-grpo-nostd-gs4-tp1-tk0-pt80000-lr1e-6-bs600-seed1 Updated Jul 16 • 1
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published Jul 9 • 77
Rank-Then-Act: Reward-Free Control from Frame-Order Progress Paper • 2607.01897 • Published Jul 2 • 7