ParaTempo: Efficient Parallel Reasoning via Temporal Confidence Paper • 2608.16425 • Published 21 days ago • 40
Apodex Discovery: Reality Benchmarks and Environments for Evaluating and Building Discoverative Artificial Intelligence Paper • 2608.11341 • Published 28 days ago • 66
Beyond Final Scores: A Systematic Evaluation of Agents for Long-Horizon AI Research and Development Paper • 2608.13417 • Published 26 days ago • 58
Large Discovery Models: Empirically-grounded Model-Based Open-Ended Search Paper • 2608.15669 • Published 23 days ago • 63
ASI-Bench: At the Dawn of Artificial Superintelligence Paper • 2608.17271 • Published 21 days ago • 65
ChronoVision: Temporal Reasoning via Latent State Reconstruction Paper • 2608.05631 • Published Aug 6 • 40
Apple-π: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published Jul 17 • 44
UP: Unbounded Positive Asymmetric Optimization for Breaking the Exploration-Stability Dilemma Paper • 2607.06987 • Published Jul 8 • 9
Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation Paper • 2607.08758 • Published Jul 9 • 42
Holistic Data Scheduler for LLM Pre-training via Multi-Objective Reinforcement Learning Paper • 2606.24133 • Published Jun 23 • 12
Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining Paper • 2606.16246 • Published Jun 19 • 4
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Paper • 2606.22388 • Published Jun 21 • 97
EurekAgent: Agent Environment Engineering is All You Need For Autonomous Scientific Discovery Paper • 2606.13662 • Published Jun 11 • 33
Swift Sampling: Selecting Temporal Surprises via Taylor Series Paper • 2605.22678 • Published May 21 • 11