The results show that harness-managed control flow can substantially improve the effectiveness of the smallest models and suggest that small models can be useful for repository repair when responsibilities are divided across explicit stages that can be independently assigned to the component best suited to each.
Francesco Dente, Dario Satriani, Donatello Santoro et al.· 0 citations
This work highlights that jointly satisfying functional and structural requirements remains a key open challenge for coding agents, and reveals a phenomenon of constraint decay: as structural requirements accumulate, agent performance exhibits a substantial decline.
Francesco Dente, Dario Satriani, Paolo Papotti· arXiv.org· 5 citations
Fact-checking is crucial for combating misinformation, and computational methods are essential for scalability. The most effective approaches leverage neural models that use domain-specific evidence to validate claims. However, these models often act as black boxes, providing labels without explaining the rationale beh...
Jean-Flavien Bussotti, Andrea Baraldi, Francesco Guerra et al.· IEEE Access· 0 citations
It is argued that LLM-native data systems should expose selected inference-time mechanisms to the optimizer as physical design choices and advocate constructing Pareto frontiers of candidate implementations and exposing only non-dominated choices to the optimizer.
Gabriele Sanmartino, Matthias Urban, Carsten Binnig et al.· 0 citations
SyntheticAgentTraceQA is proposed, an execution- first framework for generating scalable supervision data for tool- augmented agents and shows that execution-grounded supervision improves tool execution behavior, reference-trace agreement, and answer-generation performance on the evaluated tasks.
Hafsa Ouajdi, Francesco Giannuzzo, Alaa Boukhary et al.· arXiv.org· 1 citation· ⚡1
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.