Skip to content

Author

Preslav Nakov

We have 20 of 340 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

SQLStructEval: Structural Evaluation of LLM Text-to-SQL Generation

SQLStructEval is introduced, a framework that analyzes this behavior through canonical abstract syntax tree representations and adopts a pipeline that first generates structured intermediate representations and then deterministically compiles them into SQL, improving execution accuracy and structural agreement among co...

Yi-Xi Zhou, Fan Zhang, Zhiyu Guo et al. · 1 citation
#natural language process... Preprint Sep 2026

MIC: Explaining Image-Claim Inconsistencies in AI-Generated Multimodal Misinformation

Claims paired with AI-generated images are a rapidly growing form of misinformation. Existing automated fact-checking (AFC) methods mainly treat this as a provenance problem, detecting low-level synthesis artifacts to decide whether an image is AI-generated. However, such methods do not verify what human fact-checkers...

Rui-Hong Zeng, Jonathan Tonglet, Preslav Nakov et al. · 0 citations
#natural language process... Preprint Sep 2026

Same Scores, Different Decisions: Evaluating JEV and Language Models for Legal Document Understanding

Jev has the lowest cost and median response time among the evaluated configurations, while hosted language models achieve higher baseline accuracy and Rankings by baseline accuracy differ from rankings by correctness across every condition and repeat, although small differences in the latter do not establish a general...

Fan Zhang, Yan-Kai Chen, Zhuo-Han Xie et al. · 6 citations · ⚡1

Co-FactChecker: A Framework for Human-AI Collaborative Claim Verification Using Large Reasoning Models

Co-FactChecker is proposed, a framework for human-AI collaborative claim verification that translates expert feedback into trace-edits that introduce targeted modifications to the trace, sidestepping the shortcomings of dialogue-based interaction.

Dhruv Sahnan, Subhabrata Dutta, Tanmoy Chakraborty et al. · 1 citation

MemeLens: Multilingual Multitask VLMs for Memes

It is suggested that robust meme understanding requires multimodal training, varies substantially across semantic categories, and remains sensitive to over-specialization when models are fine-tuned on individual datasets rather than trained in a unified setting.

Ali Ezzat Shahroor, Mohamed Bayan Kmainasi, A. Hasnat et al. · 5 citations
Open access Aug 2026

Detecting AI-Generated Bulgarian Text: A Two-Step Multi-Class Classification Approach

This paper introduces a novel two-step multi-class classification system to identify varying degrees of machine involvement in Bulgarian text. As Large Language Models (LLMs) proliferate, distinguishing original human writing from machine-assisted or machine-generated content is crucial to prevent misinformation and pr...

Boyan Bogdanov, D. Georgiev, D. Dimitrov et al. · 0 citations

Can Dialects Be Steered Like Languages? Sparse Neurons and Distributed Directions in Arabic LLMs

Neuron steering reinforces dialect in some varieties when the prompt is already dialectal but cannot induce it from MSA prompts, whereas vector steering succeeds in both settings, and Arabic dialects are therefore steerable mainly through distributed rather than localized representations.

K. Elozeiri, Mervat T. Abassy, Omar Kallas et al. · 0 citations
Preprint Jul 2026

Jais 2: A Family of Arabic-Centric Open Large Language Models

Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report. The family includes, to our knowledge, the largest...

Mohamed Anwar, A. Freihat, George Ibrahim et al. · 3 citations
#natural language process... Preprint Sep 2026

How Correct Is Your Answer? A Semantic Correctness Framework for Open QA Evaluation

A semantic correctness taxonomy is introduced that assigns open-ended answers to eight ordered classes, separating verbose-but-correct answers from those contaminated by hallucinated content and CAP (Context-Aware Precision), a reference-based metric that scores question-conditioned statements using bidirectional NLI.

Elitsa Yotkova, Violeta Kastreva, Petar Velkov et al. · 0 citations
#artificial intelligence Preprint Sep 2026

EDRAC: Benchmarking Arabic Dialect Reading Comprehension

This work introduces EDRAC, the first large-scale benchmark for dialectal Arabic machine reading comprehension (MRC) and generative QA, covering five major dialects: Egyptian, Moroccan, Emirati, Syrian, and Saudi Arabic, and benchmarks Arabic-centric and multilingual LLMs on EDRAC using lexical and semantic metrics.

Noor Abo Mokh, K. Chirkunov, Teresa Lynn et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.