Aug 2026· Entropy· Vol 28, pp. 873· 0 citations· 47 references
Medicine
TL;DR
Gamma-Moment Equalization Initialization (GME-Init) is proposed, a data-aware asymmetric LoRA initialization method based on output-moment calibration that consistently improves standard LoRA and selected LoRA-style methods across the evaluated text understanding and multimodal tasks.
Abstract
Low-Rank Adaptation (LoRA) is a representative parameter-efficient fine-tuning method that reduces computational and memory costs without modifying the model architecture. Standard LoRA initializes matrix A from a symmetric distribution, such as Gaussian or Kaiming initialization, and matrix B to zero. Although this provides a statistically neutral starting point, it ignores the influence of task-specific input features on initialization. We propose Gamma-Moment Equalization Initialization (GME-Init), a data-aware asymmetric LoRA initialization method based on output-moment calibration. Using a small calibration set, GME-Init estimates the variance and skewness of target-layer outputs and adjusts the layer-wise initialization scale and asymmetry of LoRA weights, improving their statistical alignment with task-specific skewed representations. GME-Init operates only during initialization and does not change the LoRA architecture, trainable parameter count, training budget, or inference cost. We evaluate it on a GLUE subset with RoBERTa-base, integrate it with AdaLoRA and DoRA, and test it on VRSBench-VQA, VRSBench-Caption, and UCM-Caption using Qwen2.5-VL-3B-Instruct. Results show that GME-Init serves as a simple plug-in PEFT initialization module with no additional inference cost and that it consistently improves standard LoRA and selected LoRA-style methods across the evaluated text understanding and multimodal tasks.
Low-rank adaptation (LoRA) has become the standard for parameter-efficient fine-tuning of large language models. Most LoRA variants follow a uniform-LR convention, applying a single global learning rate across every rank-one component of every adapter. We show that this convention overlooks substantial within-module he...
Hui-Yi Wang, Daijiao Liu, Lina Yao et al.· 0 citations
Low-Rank Adaptation (LoRA) has become a de facto standard for parameter-efficient fine-tuning (PEFT), yet its performance is highly sensitive to initialization due to the information bottleneck imposed by low-rank decomposition. Existing approaches attempt to construct high-quality LoRA initializations by exploiting pr...
Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT), yet a performance gap can remain relative to full fine-tuning (FFT). Many LoRA variants improve the initialization or optimization of low-rank factors. At each training step, however, their first-order weight-space direction...
Yi Ouyang, Shi-Wei Li, Hao-Zhao Wang et al.· 0 citations
The proposed approach provides an efficient fine-tuning framework for applying LLMs to domain-specific engineering tasks with reduced computational overhead and a hybrid adaptation training strategy is designed using hierarchical learning rates and gradient clipping mechanisms across different modules.
This paper introduces a lightweight probe for multi-step gradients of pretrained weights that incurs no additional GPU memory cost and only marginal time overhead, and employs a spectrum-aware, importance-based rank allocation and optimal initialization derived from multi-step gradients.
Low-Rank Adaptation (LoRA) is an effective approach for adapting large pretrained models by learning low-rank weight updates. In practice, the LoRA rank is used to control an adapter's parameter budget and representational capacity. We show that this view is incomplete: while the nominal rank determines the representat...
Zi-Han Zhu, Zhe-Hang Du, Xu-Yang Chen et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.