Skip to content

Author

Ruizhe Li

We have 3 of 10 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#natural language process... Preprint Sep 2026

When Does Correction Become Repair? Mechanistic Auditing of Internal Interventions in Tool-Using LLMs

Before invoking external tools, an agentic LLM must select among a K-way action space: executing a call, seeking clarification, answering directly, or declining. While internal activation steering can alter these pre-execution decisions, conventional aggregate metrics obscure where altered states land and what collater...

Jia-Yi Li, Rui-Zhe Li · 0 citations
#machine learning Preprint Sep 2026

See it, Say it, Sorted: Mechanistic Diagnosis and Parameter-Space Mitigation of Emergent Misalignment in LLMs

Safety-aligned LLMs can exhibit emergent misalignment (EM): narrow domain adaptation unexpectedly triggers catastrophic safety failures across unrelated domains. Prior static analyses leave training dynamics unmapped, while existing defenses rely on heuristics that degrade utility. We present a dynamic, second-order ge...

Wei-Qiao Que, Rui-Zhe Li, Cheng-Yu Wang et al. · 0 citations
Preprint Jul 2026

When Uncertainty Isn't Enough: An Empirical Study of Self-Correction in Code Generation

It is suggested that cheap uncertainty estimators are insufficient on their own to improve code correctness, and that their practical value lies in serving as gating signals for costlier execution-based correction loops rather than as standalone substitutes for verification.

Pranav Rakasi, Maanas Lalwani, A. Srivastava et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.