Skip to content
Open access

GME-Init: Gamma-Moment Equalization for LoRA Initialization in Parameter-Efficient Fine-Tuning

Aug 2026 · Entropy · Vol 28, pp. 873 · 0 citations · 47 references
Medicine

TL;DR

Gamma-Moment Equalization Initialization (GME-Init) is proposed, a data-aware asymmetric LoRA initialization method based on output-moment calibration that consistently improves standard LoRA and selected LoRA-style methods across the evaluated text understanding and multimodal tasks.

Abstract

Low-Rank Adaptation (LoRA) is a representative parameter-efficient fine-tuning method that reduces computational and memory costs without modifying the model architecture. Standard LoRA initializes matrix A from a symmetric distribution, such as Gaussian or Kaiming initialization, and matrix B to zero. Although this provides a statistically neutral starting point, it ignores the influence of task-specific input features on initialization. We propose Gamma-Moment Equalization Initialization (GME-Init), a data-aware asymmetric LoRA initialization method based on output-moment calibration. Using a small calibration set, GME-Init estimates the variance and skewness of target-layer outputs and adjusts the layer-wise initialization scale and asymmetry of LoRA weights, improving their statistical alignment with task-specific skewed representations. GME-Init operates only during initialization and does not change the LoRA architecture, trainable parameter count, training budget, or inference cost. We evaluate it on a GLUE subset with RoBERTa-base, integrate it with AdaLoRA and DoRA, and test it on VRSBench-VQA, VRSBench-Caption, and UCM-Caption using Qwen2.5-VL-3B-Instruct. Results show that GME-Init serves as a simple plug-in PEFT initialization module with no additional inference cost and that it consistently improves standard LoRA and selected LoRA-style methods across the evaluated text understanding and multimodal tasks.

Read PDF

Similar papers

#machine learning Preprint Sep 2026

One Rate Is Not Enough: Adaptive Anisotropic Learning Rates for LoRA Fine-Tuning

Low-rank adaptation (LoRA) has become the standard for parameter-efficient fine-tuning of large language models. Most LoRA variants follow a uniform-LR convention, applying a single global learning rate across every rank-one component of every adapter. We show that this convention overlooks substantial within-module he...

Hui-Yi Wang, Daijiao Liu, Lina Yao et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TaRA: Training-Aware Low-Rank Adaptation Initialization

Low-Rank Adaptation (LoRA) has become a de facto standard for parameter-efficient fine-tuning (PEFT), yet its performance is highly sensitive to initialization due to the information bottleneck imposed by low-rank decomposition. Existing approaches attempt to construct high-quality LoRA initializations by exploiting pr...

Taehyeon Kim, Eunhyeok Park · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond Low-Rank Parameterization: Narrowing the Gap Between LoRA and Full Fine-Tuning via Gradient Decomposition

Low-Rank Adaptation (LoRA) is a widely used approach to parameter-efficient fine-tuning (PEFT), yet a performance gap can remain relative to full fine-tuning (FFT). Many LoRA variants improve the initialization or optimization of low-rank factors. At each training step, however, their first-order weight-space direction...

Yi Ouyang, Shi-Wei Li, Hao-Zhao Wang et al. · 0 citations
Open access Aug 2026

Research on Efficient Fine-Tuning of Large Model Parameters Using a Hybrid LoRA and IA3 Adaptation Approach

The proposed approach provides an efficient fine-tuning framework for applying LLMs to domain-specific engineering tasks with reduced computational overhead and a hybrid adaptation training strategy is designed using hierarchical learning rates and gradient clipping mechanisms across different modules.

J.-X. Bai · 0 citations
#artificial intelligence Preprint Aug 2026

LoRA-GA$^2$: Low Rank Adaptation with Multi-step Gradient Adaptive Alignment

This paper introduces a lightweight probe for multi-step gradients of pretrained weights that incurs no additional GPU memory cost and only marginal time overhead, and employs a spectrum-aware, importance-based rank allocation and optimal initialization derived from multi-step gradients.

Hao-Nan He, Xin Fan · 0 citations
#machine learning Preprint Sep 2026

Rank-Efficient LoRA via Joint Tangent-Space Optimization under Isotropic Curvature

Low-Rank Adaptation (LoRA) is an effective approach for adapting large pretrained models by learning low-rank weight updates. In practice, the LoRA rank is used to control an adapter's parameter budget and representational capacity. We show that this view is incomplete: while the nominal rank determines the representat...

Zi-Han Zhu, Zhe-Hang Du, Xu-Yang Chen et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.