Researchers increasingly use ChatGPT to revise their papers, and recent GPT versions often narrow or even retract the authors'claims. We call such changes defensive writing when the given material does not support them, and we test two explanations: the model corrects the authors'overclaiming, or it writes for an antic...
Results show that accuracy can hide candidate evidence-use failures and motivate role-aware audits for medical LLM evaluation, and show that many evidence interactions are clinically plausible rather than failures.
Using three support-annotated multi-hop QA benchmarks, this work compares matched adapted readers trained with raw context, retrieval windows, and gold-support diagnostic renderings to distinguish support-availability failures from remaining reader-interface effects.
Vision-language models can exhibit visual concept-conditioned divergence: given images containing demographic features, corporate logos, or ideological symbols, some models produce unusually uniform responses that differ from what peer models say about the same input. These behaviors evade text-only audits because visu...
This work introduces Code Monitor Red Teaming, a monitor-red-teaming protocol that fixes a public-check information boundary while varying generator pressure, verifier scaffolding, and weak-to-strong capability.
Jun-Hui Liao, Jiawen Deng, Fuji Ren et al.· 0 citations
A target-specific authorization audit is introduced that labels context factors separately for each tool and argument target and holds the task, proposition, position, and policy fixed while changing only the proposition's source authority.
Jun-Hui Liao· arXiv.org· 7 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.