Skip to content
Open access

GCCLA: Graph-Conditioned Cross-Lingual Adaptation of Large Language Models Under Extreme Data Scarcity (A Case Study in Tigrigna)

2026 · Proceedings of the 4th Workshop on Cross-Cultural Considerations in NLP (C3NLP 2026) · pp. 50-66 · 0 citations · 44 references

TL;DR

GCCLA is introduced, a graph-conditioned cross-lingual adaptation framework that integrates multilingual knowledge graphs into parameter-efficient LLM adaptation, conditioning a frozen multilingual LLM on structured semantic and typological relations to provide a strong inductive bias for data-efficient transfer.

Abstract

Adapting large language models (LLMs) to extremely low-resource languages remains challenging due to severe data scarcity and the lack of structured linguistic supervision. We introduce GCCLA , a graph-conditioned cross-lingual adaptation framework that integrates multilingual knowledge graphs into parameter-efficient LLM adaptation, conditioning a frozen multilingual LLM on structured semantic and typological relations to provide a strong inductive bias for data-efficient transfer. We instantiate and evaluate the framework through a focused case study on English-to-Amharic-to-Tigrinya transfer, where labeled data is extremely limited. By separating knowledge representation from language modeling, GCCLA stabilizes learning and improves sample efficiency in few-shot regimes. We evaluate the approach on five tasks, sentiment analysis, named entity recognition, nat-ural language inference, question answering, and extractive summarization, under extreme data scarcity, with as few as 0–1000 labeled Tigrinya examples. Experimental results show that GCCLA consistently outperforms multilingual, translation-based, and parameter-efficient baselines, achieves competitive performance with as few as 100 labeled examples, and degrades gracefully under partial graph coverage. These findings demonstrate that graph conditioning is an effective principle for data-efficient cross-lingual adaptation of LLMs advancing equitable NLP.

Read PDF

Similar papers

Open access Aug 2026

Optimizing sample selection for large language model-based entity matching using AssistEM

AssistEM, a framework for efficient LLM adaptation to EM via principled data selection, demonstrates that selective fine-tuning not only accelerates adaptation but also improves training efficiency (requiring fewer GPU hours), enabling open-source LLMs to rival–and in some cases outperform–closed-source models.

John Bosco Mugeni, Steven J. Lynden, Toshiyuki Amagasa et al. · 0 citations
Open access Jul 2026

Evolution and Adaptation of Large Language Models for Bahasa Indonesia

This paper investigates the engineering methodologies of cross-lingual vocabulary adaptation, parameter initialization heuristics, and language-adaptive pre-training strategies designed to address text overfragmentation, representational misalignment, and tokenization cost inefficiencies in Bahasa Indonesia and its low-resource regional dialects.

A. D. Alexander, S. Setiawati · 0 citations
Preprint Aug 2026

Cross-lingual Representation Learning via Centroid Intervention Fusion

Centroid Intervention Fusion is proposed, a projection fusion framework that consolidates multiple multilingual intervention projections into a single language-shared operator and outperforms the strongest prior pairwise intervention baseline by up to +3.3% across four model backbones.

Wei Sun, Marie-Francine Moens · 0 citations
Preprint Aug 2026

Native Multilingual Chain-of-Thought Reasoning in Low-Resource Southeast Asian Languages

The Onramp-Sequence Cross-Distillation (OSCD) is introduced, a post-training algorithm that projects high-resource reasoning trajectories into low-resource vocabulary subspaces during generative training rollouts via an integrated translator agentic loop, ensuring the stable and efficient translation of dynamically generated reference samples for fine-tuning.

Sean Gip Lim, William-Chandra Tjhi, Hai Leong Chieu · 0 citations
Book Open access Jul 2026

Graph-Enhanced Sentence Retrieval for Multi-Document Summarization in Low-Resource Languages

This approach combines language-adaptive mixture-of-experts embeddings with graph neural networks that model discourse structure, addressing linguistic challenges across typologically diverse low-resource languages.

Xuan-Hung Le, Thi Toan Do, Hoang-Quynh Le · 0 citations
Jul 2026

LLM-Assisted Sentiment Analysis for Indonesia's Coretax Policy: A Knowledge Distillation and Human-in-the-Loop Approach

The lack of high-quality labeled datasets remains a major challenge for sentiment analysis in low-resource languages such as Indonesian, particularly in specialized domains like fiscal policy. This study investigates the effectiveness of Large Language Models (LLMs) as automated annotators within a teacher-student knowledge distillation framework. Using social media data from X related to Indonesia's Coretax system, three training scenarios were evaluated: AI-labeled data, human-labeled data, and a hybrid approach. The results show that GPT-4o achieves substantial agreement with human annotators, with a Cohen's Kappa score of 0.61. Furthermore, the student model IndoBERT trained on the combined dataset outperforms other configurations, achieving a Macro F1-score of 0.64 and a Macro ROC-AUC of 0.84. These findings indicate that while LLMs cannot fully replace human judgment, they significantly enhance scalability and enable near real-time policy evaluation in low-resource settings through effective human-AI collaboration.

Novialdi Ashari, Ulfah Oktarida Sihaloho, Novi Aulia Sari · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.