GRZO: Group-Relative Zeroth-Order Optimization for Large Language Model Fine-Tuning
GRZO is a Group-Relative Zeroth-Order optimizer that draws one pseudo-independent perturbation per mini-batch example and aggregates the per-example losses through group-relative normalization, raising the effective gradient-direction count from one to the batch size at no additional forward cost while preserving inference-level memory.