A cross-domain AI-generated IQA via content-distortion awareness (CDAQA) is proposed, designed to update the existing IQA model for AGIs and achieves higher accuracy and stability in cross-domain AGIs tasks.
The findings demonstrate that AI-generated images serve as an effective diagnostic tool for robustness evaluation and failure mode analysis, and supports reliable deployment of product image quality assessment models and highlights the value of synthetic data as a structured robustness testing resource in multimedia and e-commerce systems.
Imad Tbaileh, Huthaifa I. Ashqar· Multimedia Systems· 0 citations
Image Quality Assessment is a trending area that leads to numerous applications in computer vision and systems. However, the methods in FR-IQA suffer from biases and content. The existing methods, like CNN, have limitations in a few aspects. This paper describes the Full Reference Image Quality assessment method through a Hybrid Conformer-based Full-Reference known as Frozen Vision Transformer backbone with distortion patterns and explicit modelling. The two challenges in IQA addressed in the proposed method are multiple degradations and content separation from distortion. The proposed ViT-B/16 is frozen to keep visual representations. The multiscale features are extracted and fed to parallel modules. The model learns a debiased SVD-guided subspace. This model is trained on KADID-10K and compared over three standard datasets, CSIQ, LIVE, and TID-2013, exhibits consistent performance with SRCC values 0.912, 0.924 and 0.908, respectively. The experimental study confirms that joint degradation and SVD alignment exhibit good performance.
Bheemanaboina Mahesh, Narsaiah Domala· 2026 International Conferenc...· 0 citations
This model leverages different layers of the large language model (LLM) decoder, treating them as multiple detectors that perform synchronous distortion detection using multi-level features, which helps mitigate the ambiguous foreground-background separation commonly encountered in the S2A problem.
Ziheng Jia, Yingji Liang, Jiaying Qian et al.· 0 citations
Background: The rapid growth of digital media has increased the demand for automated image aesthetic assessment (IAA) in social media, creative industries, and e-commerce platforms. Although Vision Transformer (ViT)-based models have demonstrated promising performance, most existing approaches rely primarily on spatial representations while overlooking frequency-domain information, which captures complementary characteristics such as sharpness, texture, noise patterns, and bokeh effects.
Objective: This study proposes a Dual-Branch Late Fusion FFT-ViT architecture with concatenation-based fusion as an approach for integrating dual-domain representations to improve photo aesthetic assessment (PAA).
Methods: The proposed architecture consists of two parallel branches. The RGB branch employs a Vision Transformer (ViT-Small) to extract spatial and compositional features, while the FFT branch utilizes ViT-Tiny to capture frequency-domain characteristics associated with texture, sharpness, and image details. The extracted features from both branches are fused using a concatenation strategy before being passed to the regression layer.
Results: The proposed Dual-Branch FFT-ViT with concatenation fusion achieved the best performance, obtaining a PLCC of 0.7336, SRCC of 0.7347, MSE of 0.0193, MAE of 0.1117, and RMSE of 0.1388. Compared with the RGB-only ViT baseline, the proposed model improved the PLCC score by 0.0124, demonstrating the effectiveness of integrating spatial and frequency-domain features for aesthetic score prediction.
Conclusion: This study demonstrates that integrating spatial and frequency-domain representations through a dual-branch Vision Transformer architecture enhances photo aesthetic assessment performance.
An A2I model, AudioCanvas, fine-tuned on the A2I-Set is proposed, a unified, high-quality tri-modal dataset specifically designed for audio-visual research, including audio-conditioned image generation.
Dongxu Ge, Shansong Liu, Cheng Gong et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.