Skip to content
Open access

Automatic Gujarati Text Summarization Using Natural Language Processing: A Gujarati-Specific Abstractive Framework

Aug 2026 · International journal of computer information systems and industrial management applications · Vol 18, pp. 333-341 · 0 citations

TL;DR

A Gujarati specific abstractive text summarization system that leverages an improved multilingual mT5 model with a linguistic preprocessing step on the input text, morphological normalization, named entity preservation and coverage-aware decoding step to achieve better quality of summarization, grammatical correctness, semantic consistency, computational efficiency, and solve the issues with low-resource languages from India.

Abstract

With the ever-increasing digital information, there is a high demand for automatic text summarization systems to produce meaningful summaries from longer documents. Despite significant efforts in automatic text summarization for high resource languages like English, research on automatic text summarization in Gujarati is limited because of the lack of linguistic resources, a lack of annotated datasets, and the difficult grammar of the Gujarati language. Previous multilingual transformer models like mBART, IndicBART and mT5 have shown promising results; however, they tend to produce grammatically incorrect, repetitive and incongruent summaries for Gujarati documents. In this study, we present a Gujarati specific abstractive text summarization system that leverages an improved multilingual mT5 model with a linguistic preprocessing step on the input text, morphological normalization, named entity preservation and coverage-aware decoding step. The proposed system takes as input any paragraph or article of Gujarati text or any document of large text, and produces one or two sentence summaries which are semantically equivalent to the original document and contain the information. The proposed framework would achieve better quality of summarization, grammatical correctness, semantic consistency, computational efficiency, and solve the issues with low-resource languages from India. To prove the superiority of the proposed model over the existing multilingual summarization models, automatic evaluation metrics such as ROUGE, BLEU, BERTScore will be employed in addition to human evaluation.

Read PDF

Similar papers

Open access Aug 2026

AI-Driven Kannada Document Summarization Using Optical Character Recognition and Natural Language Processing: A Web-Based Implementation Framework

The developed application is an example of the effective use of Artificial Intelligence in processing documents in regional languages and lays the groundwork for creating automated systems for document management.

Apoorva S., Usha B. S., S. Darshan · 0 citations
Open access Sep 2026

Semantic-Aware Hybrid Text Summarization Using Supervised Sentence Scoring and Redundancy Control

The rapid growth of digital textual content has intensified the need for automatic text summarization sys- tems that are both effective and reliable. While extractive summarization methods are interpretable and preserve factual content, they often suffer from redundancy and limited coherence. In contrast, abstractive a...

Khaoula Belila, Nedjoua Houda Kholladi, Mohammed Bedida et al. · 0 citations
Conference Jul 2026

Bharat Sum: OCR-Enabled Multilingual News Summarization and Bias Analysis Framework

India is home to an incredible number of languages which leads to the production of substantial amounts of news articles provided in their regional languages including Telugu, Tamil, Hindi, Bengali, Kannada, and Malayalam. Unfortunately, current systems for processing this data do so independently, rather than as part...

Farooq Sunar Mohammad, E.Sneha, B.Kavya et al. · 0 citations
Open access 2026

Evaluating Sentence Splitting on 19th-Century Italian Novels: a Comparative Analysis Across Approaches

The automatic segmentation of raw text into individual sentences, known as sentence splitting or sentence segmentation, is a fundamental task in text processing. Although it is often considered to be solved in standard domains such as news articles and Wikipedia pages, the performance of the system can vary significant...

Arianna Redaelli, Rachele Sprugnoli · 0 citations
#natural language process... Preprint Sep 2026

A Grapheme-Aware Indic Tokenizer for Tamil: Large-Scale Training and Intrinsic Evaluation

Tokenization forms the foundation of modern Natural Language Processing (NLP) systems by transforming raw text into discrete units that neural language models can process. The effectiveness of this process directly influences vocabulary efficiency, sequence length, computational cost, and downstream model performance....

V. HariKrishnanK, Sudarsun Santhiappan · 0 citations
Preprint Aug 2026

VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation

This work formalizes joint speech summarization and translation (JSumT), the generation of a succinct, faithful target-language summary directly from a long spoken document in a source language, and establishes a foundation for developing and evaluating multilingual systems capable of jointly interpreting, compressing,...

Yejin Jeon, Marie Maltais, Virginia Ceccatelli et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.