Hybrid Spiking LoRA: Asymmetric Bit-Width Design for Language Model Adaptation With Theoretical Neuromorphic Efficiency Potential—A Case Study on Korean NLU
This paper proposes Hybrid Spiking LoRA, an asymmetric spiking adapter that replaces the LoRA down-projection with a one-bit spiking encoder while retaining a full-precision up-projection, and evaluates the method on Korean natural language understanding tasks covering topic classification, relation extraction, natural language inference, and extractive question answering.
Abstract
Parameter-efficient fine-tuning adapts pre-trained language models to downstream tasks with reduced training cost, but widely used methods such as Low-Rank Adaptation (LoRA) still rely on conventional multiply-accumulate computation in the adapter path. This paper proposes Hybrid Spiking LoRA, an asymmetric spiking adapter that replaces the LoRA down-projection with a one-bit spiking encoder while retaining a full-precision up-projection. The design introduces sparse event-driven computation into the low-rank adapter while preserving continuous reconstruction capacity for language model adaptation. We evaluate the method on Korean natural language understanding tasks covering topic classification, relation extraction, natural language inference, and extractive question answering. Across timestep, rank, membrane time constant, and bit-width ablations, the encoder-only spiking design consistently outperforms fully spiking alternatives and remains competitive with standard LoRA. Hybrid Spiking LoRA retains 95.9% of standard LoRA performance on KLUE topic classification at rank 16 and 96.8% in a matched five-seed KorQuAD 2.1 comparison, where Hybrid- $T{=}4$ achieves $79.92~\pm ~0.30$ F1. On relation extraction, it underperforms standard LoRA on average but clearly exceeds fully spiking variants, and single-seed bit-width gains over FP32 are interpreted as seed-dependent preliminary observations rather than a general performance claim. Operation-level analysis indicates substantial theoretical adapter-path energy-saving potential from sparse accumulate operations, although current graphics processing unit simulation introduces latency overhead. These results position asymmetric spiking adapters as a feasible neuromorphic-compatible direction for language model adaptation, while motivating future hardware-level validation.
Extensive experiments on static and neuromorphic benchmarks show that lower-bit BASC models match or outperform higher-bit baselines and retain this accuracy advantage after structured pruning, while further reducing model storage and synaptic operations.
Linliang Chen, Yan Zhong, Xin Liu et al.· 0 citations
This work analyzes flaws of conventional conversion pipelines from residual membrane potential statistics and proposes a novel conversion strategy combining dynamic initial potential tuning and feature enhancement, which generalizes to ReLU CNNs, ANN Transformers, and multi-threshold SNN variants.
Zirui Chen, Zihan Huang, Tong Bu et al.· 0 citations
Adaptive Fission is proposed, a post-training encoding technique that selectively splits high-sensitivity neurons into groups with varying scales and weights that enables neuron-specific, on-demand precision and threshold allocation while introducing minimal spatial overhead.
Yizhou Jiang, Feng Chen, Yihan Li et al.· Neural Information Processin...· 2 citations
PTQ4SNN is proposed, a membrane-aware post-training quantization framework that jointly quantizes weights and recurrent membrane states using only a small calibration set and effectively preserves model accuracy under W4 quantization and approximately 4-bit membrane precision.
Hui Xie, Tong Shi, Haotong Qin et al.· 0 citations
It is argued the energy dividend of sparsity is not a property of SNNs but of the task, and the ceiling is formalized with an information-theoretic bound and confirmed: the floor rises with memory load, falls with state width, and (refuting a naive memory-only reading) rises with task difficulty.
Spiking Neural Networks (SNNs) enable event-driven computation with sparse activations, but building multimodal Transformers on SNNs is hindered by unstable training in deep spiking stacks and the mismatch between dense softmax attention and spike-based communication. We propose SMM Transformer, an SNN-based multimodal Transformer framework that combines (i)PLMP, a Parallel LIF with Multistage Learnable Parameters neuron and a tailored P-STBP algorithm for stable deep SNN training, (ii) SMSA, an attention-inspired spike-driven token-mixing module that replaces dense pairwise softmax attention with channel-wise spike co-activation and self-compensation, and (iii)SMoE, a spiking mixture-of-experts module for modality-aware fusion. Across visual and multimodal benchmarks, SMM Transformer achieves competitive accuracy compared to ANN baselines. Under a standard MAC/AC arithmetic model, SMSA reduces the estimated operator-level compute energy of the attention module by up to 97%, while whole-model profiling shows more moderate but consistent efficiency gains.
Xiubo Liang, Jinxing Han, Yuke Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.