Skip to content

Cross-Domain AI-Generated Image Quality Assessment via Content-Distortion Awareness

· 0 citations · 39 references

TL;DR

A cross-domain AI-generated IQA via content-distortion awareness (CDAQA) is proposed, designed to update the existing IQA model for AGIs and achieves higher accuracy and stability in cross-domain AGIs tasks.

View source

Similar papers

Jul 2026

Evaluating domain generalization of product image quality assessment models using AI-generated images

The findings demonstrate that AI-generated images serve as an effective diagnostic tool for robustness evaluation and failure mode analysis, and supports reliable deployment of product image quality assessment models and highlights the value of synthetic data as a structured robustness testing resource in multimedia and e-commerce systems.

Imad Tbaileh, Huthaifa I. Ashqar · 0 citations
Conference Jul 2026

A Hybrid Conformer-based Full-Reference Image Quality Assessment for Perceptual Image Quality Analysis

Image Quality Assessment is a trending area that leads to numerous applications in computer vision and systems. However, the methods in FR-IQA suffer from biases and content. The existing methods, like CNN, have limitations in a few aspects. This paper describes the Full Reference Image Quality assessment method through a Hybrid Conformer-based Full-Reference known as Frozen Vision Transformer backbone with distortion patterns and explicit modelling. The two challenges in IQA addressed in the proposed method are multiple degradations and content separation from distortion. The proposed ViT-B/16 is frozen to keep visual representations. The multiscale features are extracted and fed to parallel modules. The model learns a debiased SVD-guided subspace. This model is trained on KADID-10K and compared over three standard datasets, CSIQ, LIVE, and TID-2013, exhibits consistent performance with SRCC values 0.912, 0.924 and 0.908, respectively. The experimental study confirms that joint degradation and SVD alignment exhibit good performance.

Bheemanaboina Mahesh, Narsaiah Domala · 0 citations
Preprint Aug 2026

Visual Distortion Detection in UGC Images Using Large Multimodal Models

This model leverages different layers of the large language model (LLM) decoder, treating them as multiple detectors that perform synchronous distortion detection using multi-level features, which helps mitigate the ambiguous foreground-background separation commonly encountered in the S2A problem.

Ziheng Jia, Yingji Liang, Jiaying Qian et al. · 0 citations
Open access Aug 2026

Photo Aesthetic Assessment via Spatial and Frequency Domain Fusion Using a Vision Transformer Approach

Background: The rapid growth of digital media has increased the demand for automated image aesthetic assessment (IAA) in social media, creative industries, and e-commerce platforms. Although Vision Transformer (ViT)-based models have demonstrated promising performance, most existing approaches rely primarily on spatial representations while overlooking frequency-domain information, which captures complementary characteristics such as sharpness, texture, noise patterns, and bokeh effects. Objective: This study proposes a Dual-Branch Late Fusion FFT-ViT architecture with concatenation-based fusion as an approach for integrating dual-domain representations to improve photo aesthetic assessment (PAA). Methods: The proposed architecture consists of two parallel branches. The RGB branch employs a Vision Transformer (ViT-Small) to extract spatial and compositional features, while the FFT branch utilizes ViT-Tiny to capture frequency-domain characteristics associated with texture, sharpness, and image details. The extracted features from both branches are fused using a concatenation strategy before being passed to the regression layer. Results: The proposed Dual-Branch FFT-ViT with concatenation fusion achieved the best performance, obtaining a PLCC of 0.7336, SRCC of 0.7347, MSE of 0.0193, MAE of 0.1117, and RMSE of 0.1388. Compared with the RGB-only ViT baseline, the proposed model improved the PLCC score by 0.0124, demonstrating the effectiveness of integrating spatial and frequency-domain features for aesthetic score prediction. Conclusion: This study demonstrates that integrating spatial and frequency-domain representations through a dual-branch Vision Transformer architecture enhances photo aesthetic assessment performance.

Reza Rachmadan, Shintami Chusnul Hidayati · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.