A cumulative token cap can fall inside a mathematical derivation, forcing a test-time controller to choose between stopping at the cap (strict) and allowing the current attempt to finish (advisory). We measure this boundary choice with paired offline replays of 19,200 public traces: 120 AIME, BrUMO and HMMT problems an...
Gui-Lin Zhang, Zi-Qi Tan, Wu-Lan Guo et al.· 0 citations
It is found that progressive disclosure improves skill-retrieval quality but marginally degrades overall latency, and progressive disclosure of skills as needed improves skill-retrieval quality but marginally degrades overall latency.
Gui-Lin Zhang, Kai Zhao, Priyanka Mudgal et al.· 0 citations
Comparisons between AutoML systems at short time budgets -- tens of seconds rather than hours -- are common in tool READMEs and workshop papers, and they are easy to get wrong. We report a case study in which a simple AutoML engine, Orcetra, appeared to beat FLAML and AutoGluon on 513 OpenML datasets, winning 57.1% of...
Guilin Zhang, Kai Zhao· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.