AI-workflow governance cannot be reduced to checking the final answer: an apparently safe answer may rest on a fragile evidence path that ordinary evaluation cannot see, localize or govern, so SSEG moves governance below surface-level output checking into auditable, path-specific decisions about intervention, revalidation and release.
Abstract
AI-workflow governance cannot be reduced to checking the final answer: an apparently safe answer may rest on a fragile evidence path that ordinary evaluation cannot see, localize or govern. We call this hidden fragility a"structural iceberg": hallucinations and unsupported claims may form its visible tip, while consequential weakness remains submerged. Stochastic semantic evidence graphs (SSEGs) expose these icebergs by preserving workflow channels, propagating local uncertainty and identifying the hidden paths on which an apparently safe output depends. ALCE and RAGTruth show that visible failures at the tip -unsupported citations and hallucinated spans -rest on distinct submerged weaknesses and therefore require different interventions. Across retrieval, tool-use and controlled stress tests, SSEG localizes those weaknesses, supports targeted repair, produces no false automatic passes in 35,000 known-truth cases and reduces ToolSandbox review by 28.8% across 96 executions from two agent models. The same structural view carries into end-to-end governance: in a separately sealed 1,200-case FinGovBench study, adding SSEG to GPT-OSS-20B reduces unsafe releases from 452/660 to 8/660 while releasing all 540 safe cases and correctly distinguishing 592/600 matched workflow pairs. An unchanged-gate transfer to Qwen3-8B releases all 540 safe cases and none of 660 unsafe cases, whereas flat-UQ releases 520 unsafe cases. SSEG therefore moves governance below surface-level output checking, turning hidden evidence dependencies into auditable, path-specific decisions about intervention, revalidation and release.
This publication proposes a definition and a classification of agile software development approaches and analyses ten software development methods that can be characterized as being "agile" against the defined criterion.
P. Abrahamsson, O. Salo, Jussi Ronkainen et al.· arXiv.org· 727 citations· ⚡54
The study shows that agile practices improve both informal and formal communication, but indicates that, in larger development situations involving multiple external stakeholders, a mismatch of adequate communication mechanisms can sometimes even hinder the communication.
M. Pikkarainen, Jukka Haikara, O. Salo et al.· Empirical Software Engineeri...· 401 citations· ⚡48
The results indicate that software engineering work practices are chosen opportunistically, adapted and configured to provide value under the constrains imposed by the startup context.
Nicolò Paternoster, Carmine Giardino, M. Unterkalmsteiner et al.· Information and Software Tec...· 394 citations· ⚡54
It is proved that any global minimizer of the trajectory balance objective can define a policy that samples exactly from the target distribution, and empirically demonstrate the benefits of the trajectories balance objective for GFlowNet convergence, diversity of generated samples, and robustness to long action sequenc...
Esmeralda S. Whitammer, Moksh Jain, Emmanuel Bengio et al.· Neural Information Processin...· 302 citations· ⚡60
The new ChartNet training dataset could improve the accuracy of vision-language models that help analyze business trends or interpret scientific figures.
Large language models are increasingly integrated into systems that can retrieve information, call APIs, search databases, send messages, summarize documents, and take actions on behalf of users. These capabilities make LLMs useful, but they also increase the potential impact of LLM-related attacks. The post Prompt Injection Attacks in LLM Systems appeared first on GPT-Lab.