Skip to content

Author

J. Orduña

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Aug 2026

An Empirical Study of Reward Specification and Benchmark Reliability in GRPO-based LLM Unlearning

This work compares four reward designs that span lexical suppression, anti-refusal shaping, rubric-based broad answering, and an explicit refusal contrast, and shows that optimization success is not equivalent to behavioral unlearning: RWKU forget scores, held-out completion audits, terminal training-rollout audits, and training dynamics can point to different conclusions.

Rubén Balbastre, J. Orduña, Mariano Pérez · 0 citations