Uncertainty quantification (UQ) for large language models (LLMs) aims to provide reliable measures of predictive confidence, yet current methods are often unstable under meaning-preserving perturbations. Semantically equivalent paraphrases can induce substantial variability in predictive confidence, even for methods wi...
Jiayi Xin, Evan Qiang, Zi-Han Zhu et al.· 0 citations
Watermarking provides a principled way to authenticate text generated by large language models (LLMs). In practice, however, the final text may be mixed-source, with watermark evidence surviving at only a subset of token positions after rewriting, insertion, deletion, or paraphrasing. Although prior work has studied gl...
José H. Blanchet, T. Cai, Xiang Li et al.· 1 citation· ⚡1
This paper addresses the problem of optimally estimating the watermark proportion in mixed-source texts, and proposes efficient estimators for this class of methods, and derive minimax lower bounds for any measurable estimator based on pivotal statistics, showing that their estimators achieve these lower bounds.
Xiang Li, Garrett Wen, Weiqing He et al.· arXiv.org· 5 citations· ⚡1
This work introduces A 2 -Judger, a novel MLLM-based A gentic instantiation of A uto Judger equipped with semantic-aware retrieval and dynamic memory that significantly improves sample efficiency while maintaining reliable evaluation results.
Xuanwen Ding, Chengjun Pan, Zejun Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.