Articulating Assumptions in AI-Generated Scientific Analyses through Task Decomposition
This work introduces quantity grounded semantic differencing, a multi-agent framework for analyzing and comparing scientific programs generated by LLMs, and demonstrates that the modular task decomposition enhances both transparency and reliability relative to the previous single prompt approach.