Submitted by Mingqian Feng 22 Statistical Estimation of Adversarial Risk in Large Language Models under Best-of-N Sampling Microsoft 3
Submitted by taesiri 18 WebGym: Scaling Training Environments for Visual Web Agents with Realistic Tasks Microsoft 2
Submitted by Jue Zhang 28 DoVer: Intervention-Driven Auto Debugging for LLM Multi-Agent Systems Microsoft 4
Submitted by Xiao Liang 3 Gold-Medal-Level Olympiad Geometry Solving with Efficient Heuristic Auxiliary Constructions Microsoft 14 2
Submitted by Chaoyun Zhang 16 GUI-360: A Comprehensive Dataset and Benchmark for Computer-Using Agents Microsoft 2