The results suggest that LLM computation admits useful effective descriptions via RET: high-level, dynamically meaningful variables for interpretation, prediction, and control.
Muhammed Ustaomeroglu, Guan-Nan Qu· arXiv.org· 0 citations
TRAP, a one-sided penalty on tokenwise TRA that acts only where the target model pulls ahead of its reference, brings memorization near the level of an untrained model at little utility cost, where generic regularizers barely move and differential privacy gives up most of what fine-tuning bought.
Muhammed Ustaomeroglu, Zi-Yue Xu, Han-Shen Xiao et al.· 0 citations
Activation-Keyed Momentum (AK-Momentum), which builds direction-awareness into the momentum update rule, is proposed, and it is proved that it is a valid momentum, that it applies the input-side curvature correction without matrix inversion, and that it clears stale directions faster than EMA under both a fixed and a d...
E. Hong, Guan-Nan Qu· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.