Skip to content
Preprint

EVIL-Detect for NLPCC 2026 Shared Task 6: LLM-Generated Text Detection

Aug 2026 · 0 citations · 25 references
Computer Science

TL;DR

EVIL-Detect, a multi-signal ensemble framework with conflict-aware fusion for NLPCC 2026 Shared Task 6, improves robustness under strong out-of-distribution shifts, achieving a macro-F1 score of 0.8888 and ranking first in the official evaluation.

Abstract

The rapid development of large language models (LLMs) has increased the need for reliable detection of LLM-generated text, especially in realistic Chinese scenarios involving human-written text (HWT), LLM-generated text (LGT), and LLM-refined text (HLT). This paper presents EVIL-Detect, a multi-signal ensemble framework with conflict-aware fusion for NLPCC 2026 Shared Task 6. The system integrates edit-extent regression, zero-shot likelihood-contrast signals, lexical statistics, and conservative text rules. With calibrated decision boundaries and conflict-aware integration, our system improves robustness under strong out-of-distribution shifts, achieving a macro-F1 score of 0.8888 and ranking first in the official evaluation. Our code is available at https://github.com/bbbbhrrrr/evildetect.

View source

Similar papers

Jul 2026

Detecting LLM-Generated Tokens in Human-LLM Coauthored Text

The key idea is to smooth adjacent token scores to reduce their variability, while using an adaptive Lepski-type rule to select the bandwidth according to the local authorship structure, and the proposed method achieves favorable mean square error performance in estimating the underlying signal.

Yangjun Lu, Hongyi Zhou, F. Spill et al. · 0 citations

Information Entropy for LLM-generated Text Detection

A novel method named IED is proposed, which leverages the information gain to construct the vector of which each dimension represents the information entropy of each word, and then adopts a classifier to conduct the detection.

Xiaoquan Yi, Haozhao Wang, Jingcai Guo · 0 citations

Munich_Z@GermEval Shared Task 2025:When Prompting Is Not Enough: The Limits of Large Language Models in GermEval’s 2025 Harmful Content Detection Task

This work evaluates prompting strategies for subtask 2 of the GermEval 2025 Harmful Content Detection challenge, which involves classifying whether a tweet attacks the free democratic basic order and shows that techniques such as Chain-of-Thought, In-Context Learning or Task Decomposition outperform approaches like Task Description.

Florian Ludwig, Stefan Altmann · 1 citation
Open access 2026

Advancing Machine-generated Text Detection: A Comprehensive Evaluation of Transformer-based Models

Test set results show that Decoding-Enhanced Bert with Disentangled Attention (DeBERTa) achieves the highest macro F1 − Score of 85.48%, surpassing the previously top-ranked Multi-Task Learning (MTL) system, which attains a macro F1 of 83.07%.

Batyr Sharimbayev, S. Kadyrov · 0 citations
Book Open access Aug 2026

Cost-Aware Human-LLM Collaboration for Post-OCR Corrections in Swiss Historical Newspapers

A regression-guided routing approach that prioritizes segments by predicted CER improvement, paired with a safeguard layer that detects harmful LLM corrections and routes uncertain segments to human review, and substantially outperforms standard confidence-based approaches is introduced.

Stergios Konstantinidis, Hayman Lotfy, Michalis Vlachos · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.