Skip to content
Conference

A hybrid approach to semantic text similarity combining word embeddings and classical similarity measures

Aug 2026 · International Conference on Digital Transformation: Informatics, Economics, and Education · Vol 14303, pp. 143030F - 143030F-8 · 0 citations · 14 references
Engineering

TL;DR

It is determined that the proposed hybrid model attains higher correlation with human judgment than standalone or traditional baselines, and is an efficient and scalable solution that balances computational performance with semantic accuracy for practical tasks.

Abstract

In this study, a hybrid approach to semantic text similarity combining distributed word embeddings with classical lexical similarity measures is developed. Analyzed are the limitations of modern deep learning models, namely computational overhead and weak interpretability in resource-constrained environments. Proposed is a hybrid architecture that integrates Word2Vec distributed representations with cosine and Jaccard lexical similarity metrics. Investigated is a weighted fusion mechanism that combines vector-based semantic distances with set-theoretic token overlap for robust scoring. Developed is a three-stage processing pipeline covering text preprocessing, sentence embedding generation, and similarity computation and fusion. Established is a tunable weighting parameter that experimentally balances semantic depth against lexical matching precision. Conducted are experimental evaluations on benchmark semantic textual similarity and paraphrase detection datasets using classification metrics. Determined is that the proposed hybrid model attains higher correlation with human judgment than standalone or traditional baselines. Demonstrated is a notable reduction of error rates for exact lexical matches frequently missed by vector-only models. Presented is an efficient and scalable solution that balances computational performance with semantic accuracy for practical tasks.

View source

Similar papers

Conference Aug 2026

Evaluation of the BERT model for text semantic similarity

This study validates the effectiveness of the BERT model in semantic similarity calculation, providing more accurate technical support for related application scenarios, and laying the foundation for subsequent model optimization and lightweighting research.

Jiachen Gao · 0 citations
Jul 2026

A novel semantic–syntactic hybrid plagiarism detection system based on word embeddings and similarity measures

Experimental results indicate that the proposed approach achieves competitive performance compared to existing plagiarism detection systems, and the comparative analysis highlights the strengths and limitations of different word embedding models across datasets.

Malya Singh, Vishal Gupta · 0 citations
2026

A Multi-Similarity Neural Network for Paraphrase Detection

The results suggest that integrating diverse similarity measures with neural networks enhances the identification of both explicit and nuanced paraphrases, thereby supporting advancements in text analysis and plagiarism detection systems.

Emad Nabil · 0 citations
Conference Aug 2026

Optimized architecture based on multilayer contextual embeddings for evaluating semantic similarity of texts in the Uzbek language

Semantic text similarity is an especially hard task that can be performed in the Uzbek language because of its rich morphological and the absence of annotated data. The given paper introduces a very effective Bidirectional Encoder design using the monolingual model named “Bidirectional Encoder Representation from Transformers for Uzbek language” and optimized to provide scalable semantic search. A combination of knowledge distillation and metric learning are applied in a semi-supervised approach that is employed by the model. Triplet Loss with Hard Negative Mining is used to enhance the discriminative ability of the vector space. One such innovation is supporting the Matryoshka Representation Learning, which allows the model to produce dynamically truncated dimensions of embeddings. The architecture proposed wound up with a Spearman correlation of 0.835 within the Uzbek test set, topping the state-of-the-art results. Furthermore, Matryoshka Representation Learning achieves 6-fold compression of vectors with a negligible accuracy drop (ρ=0.816), having high-computation efficiency to implement Natural Language Processing systems in low-resource systems.

B. Muminov, N. Allaberganova, Olimjon Mamadiyorov et al. · 0 citations
Open access Aug 2026

Use BiLSTM with Attention Mechanism to Optimize the Accuracy of Word Meaning Correspondence in Technical Texts

A robust WSD model that integrates a bidirectional long shortterm memory network (BiLSTM) with an attention mechanism, specifically designed for Chinese patent texts is proposed, providing reliable support for knowledge mining and intelligent text processing in technical domains.

L. X. Gao, T. Dong, M. H. Yang · 0 citations
Open access 2026

AraBERT-Based Semantic Lexicon Induction for Arabic Academic Text: Query Expansion, Document Clustering, and Classification

The scarcity of high-quality semantic lexical resources for Arabic academic text represents a critical bottleneck for natural language processing (NLP) applications, including query expansion, document retrieval, and information organization. This paper presents a framework for inducing semantic neighbor lexicons from the Arabic Research Papers Dataset (ARPD), a publicly available corpus of 2,011 Arabic academic documents spanning seven scientific domains. We exploit type-level representations derived from AraBERT to compute dense, L2-normalized word embeddings and apply GPU-accelerated cosine similarity search to retrieve up to three semantically induced neighbors per vocabulary item. Experiments are conducted on both the raw and preprocessed versions of ARPD. The raw-corpus lexicon covers 168,866 unique terms, and the preprocessed-corpus lexicon covers 159,364, both with 100% three-neighbor coverage and a mean top-1 cosine similarity of 0.91. For document clustering, replacing TF-IDF bag-of-words with AraBERT document embeddings raises the Silhouette coefficient from 0.070 to 0.645 on the raw corpus (+ 820.0%) and to 0.755 on the preprocessed corpus (+ 963.4%), alongside substantial reductions in the Davies–Bouldin index. A pilot expansion experiment under class-based relevance finds no retrieval gain from lexicon expansion, which we report alongside the root-level and expert analyses that explain it. For classification, TF-IDF with only light normalization achieves 99.17% accuracy, exceeding the published benchmark of 99.00% that required heavy preprocessing. Comprehensive comparisons against static embedding baselines (Word2Vec, FastText, GloVe) under both corpus conditions quantify how preprocessing depth interacts with representation type. A comparison against MARBERTv2, AraELECTRA, and CAMeLBERT-MSA under leakage-free, validation-based checkpoint selection shows MARBERTv2 attaining the highest accuracy (99.67% raw, 99.50% preprocessed); we emphasize that differences among the top-performing methods correspond to only a few documents and are not statistically significant, while semantic-neighbor augmentation yields a nominal + 1.82 pp gain for AraBERT on the preprocessed corpus (McNemar $p=0.027$ ).

Ahmad Al Smadi, Yasser Allaham, Lara Rbabah · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.