Critic-Free Pretraining for Efficient Online Reinforcement Learning Fine-Tuning
Critic-Free Pretraining is introduced: an efficient paradigm that completely abandons the approach of offline critic training, allowing a freshly initialized critic to adapt without inheriting biased estimates.