#artificial intelligence
Apr 2026
Pareto-Optimal Offline Reinforcement Learning via Smooth Tchebysheff Scalarization
STOMP is a powerful, robust multi-objective alignment algorithm that can meaningfully improve post-training in multiple domains and achieves or ties for the highest hypervolumes on 16/18 protein tasks and 5/6 natural language tasks.
Aadyot Bhatnagar, Peter Mørch Groth, Sebastian Ibarraran et al.
· arXiv.org · 0 citations