Skip to content

Information Entropy for LLM-generated Text Detection

· 0 citations · 42 references

TL;DR

A novel method named IED is proposed, which leverages the information gain to construct the vector of which each dimension represents the information entropy of each word, and then adopts a classifier to conduct the detection.

View source

Similar papers

Jul 2026

Detecting LLM-Generated Tokens in Human-LLM Coauthored Text

The key idea is to smooth adjacent token scores to reduce their variability, while using an adaptive Lepski-type rule to select the bandwidth according to the local authorship structure, and the proposed method achieves favorable mean square error performance in estimating the underlying signal.

Yangjun Lu, Hongyi Zhou, F. Spill et al. · 0 citations

VaryBalance: Detecting LLM-generated Text through Variation

The core of VaryBalance is that, compared to LLM-generated texts, there is a greater difference between human texts and their rewritten version via LLMs, and quantifies this through Mean Squared Deviation and distinguishes human texts and LLM-generated texts.

Xuecong Li, Xiaohong Li, Qiang Hu et al. · 0 citations
Preprint Aug 2026

EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection

EVIL-Detect, a multi-signal ensemble framework with conflict-aware fusion for NLPCC 2026 Shared Task 6, improves robustness under strong out-of-distribution shifts, achieving a macro-F1 score of 0.8888 and ranking first in the official evaluation.

Hongrui Bao, Hangyu Rong, Zhuo Wang et al. · 0 citations
Open access 2026

Combining probabilistic features and semantic features for AI-Generated text detection

The proliferation of Large Language Models (LLMs) such as ChatGPT and Gemini has resulted in a surge of AI-generated text across various domains. However, the widespread use of this technology raises concerns regarding the generation of misinformation and malicious content. To address this challenge, we propose a novel AI-generated Text Detection model combining Probabilistic and Semantic features (ATDPS). Our model extracts semantic features using a pre-trained language model and combines them with probabilistic features generated by multiple LLMs. A temporal convolutional network is employed to process sequence probabilistic features, effectively capturing temporal characteristics within the text. To ensure data coherence and diversity, our dataset includes text generated by a variety of LLMs, including the latest models like GPT-4. Experimental results demonstrate ATDPS's superior performance over existing baselines in terms of accuracy, precision, recall and F1 score, highlighting its potential and effectiveness in detecting AI-generated text.

Yang Yu, Wang Gao · 0 citations
Open access Aug 2026

Text Classification Using Large Language Models to Detect AI and Human-Generated Sentences

With the rapid development of the technology of Artificial Intelligence (AI), especially Large Language Models (LLMs) like ChatGPT, GPT-4 and Gemini, systems have become able to generate texts very similar to human writing. This similarity has been a boon to various sectors, but also poses fresh challenges of content authenticity and data integrity. A significant challenge is finding a way to automatically and accurately distinguish AI-generated sentences from human-written ones. This research focuses on building a text classification model that uses Large Language Models to distinguish between AI-generated and human-written sentences. This approach of research is based on recent research models that combine deep learning and classification using LLMs . The research process involves gathering human and AI-generated text data, pre-processing the text to normalize and tokenize it, feature extraction using the embeddings of large language models like IndoBERT, training the binary classification model, and assessing the model's performance using the metrics accuracy, precision, recall, and F1-score. Keywords: Text clustering, Large language models, AI text identification, Text authenticity.  

Arfan Rahman Yudiantoro, Prati Hutari Gani, Donni Richasdy · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.