Skip to content
Preprint

LHSDet: High-Resolution AI-Generated Image Detection via Visual Question Answering

Aug 2026 · 0 citations · 54 references
Computer Science Engineering

TL;DR

This work formats the AI-generated image detection task as a Visual Question Answering problem, leveraging a fine-tuned vision-language framework to fully exploit the complementary information between visual and textual modalities, and proposes a novel high-resolution AI-generated image detector, termed LHSDet.

Abstract

Driven by advances in diffusion models and autoregressive models, the fidelity and resolution of AI-generated images now rival those of real images. However, existing AI-generated image detection methods often downsample the images, inevitably overlooking critical low-level texture details in high-resolution AI-generated images, therefore limiting their detection performance. In addition, the ceaseless emergence of unknown generative models makes large-scale pre-training datasets inaccessible. To address these challenges, we propose a novel high-resolution AI-generated image detector, termed LHSDet. Specifically, we formulate the AI-generated image detection task as a Visual Question Answering problem, leveraging a fine-tuned vision-language framework to fully exploit the complementary information between visual and textual modalities. Recognizing that the default visual encoder of existing vision-language models is not tailored for AI-generated image detection, we redesign a visual encoder to better capture both the low-level and high-level artifacts inherent in AI-generated images. Furthermore, we incorporate a semantic-level textual branch to enable multi-modal feature fusion and detection. Consequently, LHSDet employs a triple-branch architecture to extract complementary multi-modal features: a low-level visual branch that aggregates non-overlapping patches for local texture cues, a high-level visual branch based on SigLIP2 for global perception feature extraction, and a semantic-level textual branch that generates captions using BLIP-2. Extensive experimental results demonstrate that LHSDet achieves high detection accuracy and robust performance across diverse generative models, including both diffusion and autoregressive models.

View source

Similar papers

Open access Jul 2026

Attention-Based Deep Learning Pipeline for AI-Created Image Recognition

The proposed Attention-Based Deep Learning Pipeline of AI-Created Image Recognition incorporates three integrated branches, including low-level statistical feature extraction, high-level semantic representation learning, and attention-based feature refinement mechanism, which support the robustness and generalization ability of the proposed model in detecting AI-generated images in a variety of generators and conditions.

Nadia Ali · 0 citations
Preprint Aug 2026

Generated Images Are Easier to Forget: A Machine Unlearning Perspective for Synthetic Image Detection

This work establishes a new paradigm for generated image detection by recasting the detection task as a problem of machine unlearning, and introduces two detection methods: data-free detection, which prunes model parameters to induce unlearning without data access, and data-driven detection, which optimizes LVMs to unlearn knowledge tied to generated images.

Jun Nie, Yonggang Zhang, Tongliang Liu et al. · 0 citations
Preprint Aug 2026

UC-VLM: Consistency-Driven Learning for AI-Generated Image Detection with Vision-Language Large Models

UC-VLM is a unified multi-stage binary-supervised framework that consistently reuses the same authenticity labels for visual adaptation and label-conditioned text generation, while leveraging automatically optimized instructions to reduce prompt sensitivity without requiring human-written rationales or hand-crafted prompts.

Lei Tan, Shuwei Li, Mohan S. Kankanhalli et al. · 0 citations
Open access Jul 2026

AI-Driven Image Synthesis from Textual Descriptions Using Stable Diffusion

An AI-based image generation system is introduced that utilizes a Stable Diffusion model fine-tuned with Low-Rank Adaptation (LoRA) for domain-specific image generation, demonstrating that SD+LoRA is an efficient and scalable domain-specific text-to-image generation system.

Ankam Pavitra, R. Mallikharjun, Dr. L Jagadeesh Naik · 0 citations

PPM-CLIP: Probabilistic Prompt Modeling for Generalizable AI-Generated Image Detection

PPM-CLIP is proposed, a new framework that shifts from static classification to conditional generative modeling based on the CLIP vision-language model, and a Probabilistic Prompt Modeling module is used as a generator that produces an adaptive distribution of prompts according to the input image.

Xinyu Wang, Yingxin Lai, Zhiming Luo et al. · 0 citations
Jul 2026

Can Vision-Language Models Reason about AI Edits in Images?

This work investigates whether VLMs can be trained to reason about AI-generated image edits using reinforcement learning (RL) rather than explicit reasoning supervision, and introduces effective intersection over union (eff-IoU), a unified metric to jointly evaluate detection and localization.

Darsha Udayanga, Pin-Yu Chen, Payel Das et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.