Conformalized Large Language Models under Configuration Shift
It is found that configuration shift consistently erodes CP validity, often driving empirical coverage below the target, and coverage lower bounds are derived that attribute this loss to a discrepancy between calibration and test score distributions.
Yuqicheng Zhu, Jia-Lin Yu, Lin Li et al.
· 0 citations