Skip to content
Preprint

A Benchmark Dataset for MLLM-Generated Image Detection: GPT Image2&Nano Banana2

Aug 2026 · 1 citation · ⚡ 1 influential · 26 references
Computer Science

TL;DR

A benchmark dataset for detecting images generated by MLLMs is constructed and a structural-artifact-prior-guided dual-stream prompt framework (SAP-DSP) is proposed as a strong baseline, using dual-stream prompt learning and structure-aware routing fusion to improve representation learning.

Abstract

The realism of images generated by multimodal large language models (MLLMs), such as GPT Image2 and Nano Banana2, has improved rapidly in recent years. Compared with early generative models, current models have made clear progress in text rendering. They can produce high-quality images that closely resemble real-world application scenarios. The enhanced generation capabilities of current MLLMs pose increasingly severe challenges to AI-generated image detection. Detection is no longer limited to identifying obvious artifacts left by early generators. Instead, it requires systematic and realistic benchmarks for the new generation of generated content. However, most existing benchmarks are still built around early generative models and cannot fully evaluate the forensic challenges introduced by high-quality and multi-form generated images. To address this gap, this paper constructs a benchmark dataset for detecting images generated by MLLMs. The benchmark covers several realistic application scenarios and adopts three generation protocols to simulate direct generation, reference-based reconstruction, and local editing. Based on this benchmark, we evaluate detector degradation from traditional scenarios to MLLM-generated images and analyze false positive rates and false negative rates across three sample types, revealing the failure modes of existing methods. We further propose a structural-artifact-prior-guided dual-stream prompt framework (SAP-DSP) as a strong baseline. SAP-DSP uses dual-stream prompt learning and structure-aware routing fusion to improve representation learning. Extensive experiments show that the proposed benchmark exposes the performance degradation of existing detectors on high-quality generated images, while SAP-DSP achieves more stable detection results on this benchmark. Our code and dataset are publicly available at https://github.com/xbrainnet/SAP-DSP.

View source

Similar papers

#artificial intelligence Preprint Sep 2026

REALIS: A Curated Dataset for Studying the Challenges of AI Image Detection

AI-generated image detectors are often evaluated on benchmarks where real and synthetic images differ in content, quality, or generation artifacts, allowing models to rely on dataset-specific cues and fail on unfamiliar generators or processed images. Existing datasets provide limited support for evaluating these chall...

A. Gushchin, Khaled Abud, G. Bychkov et al. · 0 citations
Conference Open access 2025

Comparative Analysis of Image Generation Based on GAN, VAE, and Diffusion Models

: The core task of image generation models is to generate visual content that meets specific requirements based on given inputs. The wide application of artificial intelligence has accelerated the development of generative technologies, and image generation has had a significant impact across various fields in real-wor...

Qirui Guo, Yuxing Hu, Zhe-Jia Wang · 0 citations
Preprint Sep 2026

Evaluating the Evaluators: Diagnosing Large Multimodal Models for AI-Generated Image Assessment

With the rapid advancement of text-to-image (T2I) generation, robust evaluation becomes critical yet challenging, as traditional metrics fail to capture fine-grained alignment and generative artifacts. While large multimodal models (LMMs) are increasingly adopted as evaluators, existing benchmarks typically study seman...

Yu Zhao, Jia-Rui Wang, Hui-Yu Duan et al. · 0 citations
#generative ai Open access Sep 2026

Transfer Learning-Based Detection of AI-Generated Image

This study investigates the automatic classification of real and AI-generated flower images using fine-tuned transfer learning models and shows that Swin Transformer-Tiny achieved the best overall performance, reaching an F1-score of 88.64% and outperforming the other architectures.

Mehtap Ülker · 0 citations
Conference Open access 2025

Enhanced Anime Image Generation by Comparative Analysis of DCGAN, LSGAN, and WGAN-GP

: In recent years, anime-style image generation has become a prominent direction within generative adversarial network (GAN) research. However, a systematic exploration into the performance differences among various GAN architectures, specifically for anime face generation is still lacking. Therefore, this study utiliz...

Bing-Hui He · 0 citations
Preprint Aug 2026

AGIDefect-4K: A Richly Annotated Dataset for AI-Generated Image Defect Detection, Localization and Explanation

Generative AI can now produce highly realistic images, yet current models still exhibit subtle but critical defects that undermine their reliability. While existing AI-generated image (AGI) evaluation benchmarks have made notable progress, comprehensive AGI defect diagnosis remains underexplored. To bridge this gap, we...

Xiang-Fei Sheng, Weidong Zou, Tianjiao Gu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.