Large language models are increasingly used to read earnings calls, investor-relations Q&A, guidance, and disclosure language. In this setting, supervised financial NLP benchmarks can become evidence for vendor selection, deployment approval, and model-risk records. Gold labels, however, do not make a benchmark score a...
Si-Di Chang, Pei-Ke Zhu, Yu-Xiao Chen et al.· IEEE Conference on Computati...· 0 citations
The result is a bounded rule for interpreting aggregate agent behavior: first establish exposure, then score change, and abstain when the trace cannot support the claim.
An AI evaluation can be perfectly reproducible and still support the wrong claim. This risk is acute in closed-loop systems: policy determines visited states, observable components, and which failures leave a measurable trace. We propose a claim-safe protocol with three actions. Refuse: abstain when a clean reference s...
Synthetic perturbations appear to offer inexpensive calibration data for LLM evaluators in biomedical ML, where expert review is scarce. Yet a planted mutation key is neither a detector output nor automatically human ground truth. We formalize four distinct ledgers: planted perturbations, independent detector outputs,...
ClaimReceipt, a claim-relative receipt specification and selective verifier that binds typed transaction evidence to a signed experiment manifest and returns PASS, INVALID, or INCONCLUSIVE per claim, is introduced.
The case does not show that guardrails are ineffective; it shows their apparent value is unidentified until the simulated agents and protocol pass these checks, and contributes a construct-validity contract separating incentive validity, protocol isolation, stochastic stability, and welfare accounting.
Pei-Ke Zhu, Si-Di Chang· 1 citation
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.