Earth-system analysis reconstructs changing physical processes from observations that differ in source, scale, timing, and modality. Natural hazards make this work consequential because incomplete evidence can change estimates of severity, exposure, and mechanism. We introduce EarthVerse, a benchmark that evaluates sci...
Zhiqing Cui, Xinxiang Yin, Yihong Tang et al.· 0 citations
Evaluated on diverse mathematical, algorithmic, and systems optimization tasks, CORAL sets new state-of-the-art results on 10 tasks, achieving 3-10 times higher improvement rates with far fewer evaluations than fixed evolutionary search baselines across tasks.
Ao Qu, Handi Zheng, Zi-Jian Zhou et al.· arXiv.org· 38 citations· ⚡7
The HUman-Grounded AGENT Benchmark is introduced, which rethinks human reasoning simulation along three dimensions: (i) from averaged to individualized reasoning, (ii) from behavioral mimicry to cognitive alignment, and (iii) from vignette-based to open-ended data.