OpenR1 SFT to GRPO: Token-Level Study Collection Matched OpenR1 SFT checkpoints and subsequent GRPO models for studying token-level behavior during the transition from SFT to RL. • 14 items • Updated about 18 hours ago
OpenR1 SFT to GRPO: Token-Level Study Collection Matched OpenR1 SFT checkpoints and subsequent GRPO models for studying token-level behavior during the transition from SFT to RL. • 14 items • Updated about 18 hours ago
OpenR1 SFT to GRPO: Token-Level Study Collection Matched OpenR1 SFT checkpoints and subsequent GRPO models for studying token-level behavior during the transition from SFT to RL. • 14 items • Updated about 18 hours ago
OpenR1 SFT to GRPO: Token-Level Study Collection Matched OpenR1 SFT checkpoints and subsequent GRPO models for studying token-level behavior during the transition from SFT to RL. • 14 items • Updated about 18 hours ago
OpenR1 SFT to GRPO: Token-Level Study Collection Matched OpenR1 SFT checkpoints and subsequent GRPO models for studying token-level behavior during the transition from SFT to RL. • 14 items • Updated about 18 hours ago
OpenR1 SFT to GRPO: Token-Level Study Collection Matched OpenR1 SFT checkpoints and subsequent GRPO models for studying token-level behavior during the transition from SFT to RL. • 14 items • Updated about 18 hours ago
Self-Generated Feedback Destabilizes Test-Time Training: A Causal Decomposition of Long-Horizon Adaptation Paper • 2610.05076 • Published 7 days ago • 26
OpenR1 SFT to GRPO: Token-Level Study Collection Matched OpenR1 SFT checkpoints and subsequent GRPO models for studying token-level behavior during the transition from SFT to RL. • 14 items • Updated about 18 hours ago