On-Policy Delta Distillation for Multilingual Math Reasoning Paper • 2608.05802 • Published 11 days ago • 32
HarnessOpt-Bench: Evaluating LLMs at Harness Optimization Paper • 2608.06301 • Published 11 days ago • 34
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published 14 days ago • 172
NOLLI: A Difficulty-Calibrated Puzzle Benchmark for Diagnosing the English-Korean Performance Gap Paper • 2608.04397 • Published 12 days ago • 23
Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation Paper • 2607.27372 • Published 19 days ago • 19
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis Paper • 2607.28618 • Published 18 days ago • 302
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Paper • 2607.21653 • Published 26 days ago • 32
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Paper • 2607.22529 • Published 24 days ago • 48
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Paper • 2607.20911 • Published 25 days ago • 26
Running 62 Don't Train the Model, Evolve the Harness 🌿 62 Evolving an agent's harness, not its model, on Harvey's LAB
DSWorld: A Data Science World Model for Efficient Autonomous Agents Paper • 2607.15901 • Published about 1 month ago • 12
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published Jul 16 • 106