Multi-agent debate (MAD) has emerged as a promising paradigm for improving the reasoning accuracy of large language models (LLMs) through iterative peer interaction. Communication topology plays a central role in this process, motivating increasingly sophisticated mechanisms that learn, adapt, or dynamically reconfigur...
Bo-Xuan Wang, Zhuo-Yun Li, Xiao-Wei Huang et al.· 0 citations
This paper introduces SCOPE (Sequential Conformal OOD Probing and Evaluation), a framework that selects a readable hidden layer, constructs a conformal gate with IND calibration, and uses a supermartingale e-process to certify persistent service-boundary evidence.
FragileFlow is introduced, a plug-in regularizer that uses a calibrated margin buffer to identify correct-but-fragile predictions and organize their off-class probability mass into a class-wise vulnerable-risk matrix and provides the first PAC-Bayes upper bound for this margin-aware error-flow object.
Zhuo-Yun Li, Bo-Xuan Wang, Jinwei Hu et al.· arXiv.org· 4 citations
The Reasoning Backroom is established as a general AI provenance problem whose audit requires intervention because of a pervasive provenance failure in multi-agent systems.
Jinwei Hu, Yi Qi, Xin-Miao Huang et al.· arXiv.org· 0 citations
This work proposes Logical Graph Uncertainty (LGU), a framework that explicitly models implication and incompatibility among answers, and ranks first on average among existing uncertainty measures.