Tencent Hunyuan releases ExplorationBench to test how AI systems discover rules
Original titleNew Research: We are releasing ExplorationBench, a benchmark for measuring how AI systems explore.
AISummary
Researchers from Tencent Hy, Fudan University, and Tsinghua University released ExplorationBench, a benchmark that tests whether AI systems can discover hidden rules in executable Alien World sandboxes.
Across 10 frontier systems, getting feedback from experiments outperformed thinking alone, with the best run reaching 89.0% after four rounds. The authors note that rankings barely transfer between the two worlds, and the code is listed as coming soon.
Source: Tencent Hunyuan · x.comPublished · added here