OE25𝑑𝑒𝑣, a multi-variant dataset curated from developer-written unit tests across 25 open-source Java projects spanning 56 modules, and TOGBench, an end-to-end benchmark suite for TOG, which captures six oracle categories and preserves realistic settings, are introduced.
Tasfia Tasnim, Matthew B. Dwyer, Soneya Binta Hossain· AIware· 1 citation
Future TOG systems should be evaluated not only by whether they predict the correct oracle type, but also by whether their predictions are grounded in meaningful exception-triggering evidence, to challenge the assumption that strong exception-oracle accuracy reflects robust use of exception semantics.
Soneya Binta Hossain, Matthew B. Dwyer, Tasfia Tasnim· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.