Aug 2026· IEEE transactions on consumer electronics· 0 citations· 38 references
EngineeringComputer Science
TL;DR
This work proposes MDFI (Multi-Domain Features Integration), a compressed video quality enhancement approach that features a novel Frame-Prediction Feature Transform (FPFT) module to process prediction information to enhance decoded video quality.
Abstract
The latest video coding standard, H.266/VVC, has demonstrated significant improvements in compression efficiency compared to H.265/HEVC. Despite its advanced coding techniques, H.266/VVC still faces challenges in meeting the increasing demand for higher perceptual quality and enhanced compression performance. To address these limitations, we propose MDFI (Multi-Domain Features Integration), a compressed video quality enhancement approach that features a novel Frame-Prediction Feature Transform (FPFT) module to process prediction information. Moreover, MDFI integrates a multi-domain feature fusion strategy that effectively combines spatiotemporal characteristics, cross-frequency representations, and compressed-domain prediction information to enhance decoded video quality. Additionally, we introduce a comprehensive dataset that encompasses uncompressed video sequences, corresponding reconstructed versions at multiple QP levels, and predicted frames generated from H.266/VVC compressed bitstreams, providing essential resources for developing and benchmarking video enhancement approaches. Extensive experiments demonstrate that our MDFI approach achieves superior performance to state-of-the-art methods in both objective metrics and visual quality, effectively mitigating video compression artifacts. The code is available at: https://github.com/dangdinh17/MDFI.git.
Traditional block-based video codecs, such as H.264/AVC, H.265/HEVC and H.266/VVC, rely on hand-crafted Rate-Distortion Optimization (RDO) processes that primarily minimize Mean Squared Error (MSE), which correlates poorly with human perceptual quality. While neural video compression methods can easily optimize perceptually aligned metrics like MS-SSIM, their high computational complexity limits practical deployment. This paper proposes a novel bit allocation transfer framework that bridges these two paradigms to enhance the perceptual quality of conventional video codecs. Specifically, we train a quantization step generation model using a perceptual loss within a neural video compression framework (DCVC-FM). The model takes the original frame and a motion-compensated prediction as input and outputs a quantization step map. We then derive a block-wise bit ratio from this map and convert it into a Quantization Parameter (QP) map for a traditional video codec. Experimental results on the HEVC B$\sim$D dataset demonstrate that our method achieves 20.20\%, 8.25\%, and 8.37\% bitrate savings in terms of MS-SSIM compared with the standard reference software JM-19.0, HM-16.20, and VTM-23.0, respectively, with additional gains when utilizing predicted frames. Our approach effectively transfers the implicit perceptual importance learned by neural video compression models to guide block-level bit allocation in traditional video codecs without modifying their core decoding syntax.
With the exponential growth of video traffic and the continuous evolution of video coding standards, video transcoding has become essential for existing bitstreams to benefit from the advanced features of new video compression technologies. Typically, video transcoding involves decoding an existing bitstream and re-encoding the decoded sequence into a target format. A key challenge in transcoding is the inevitable presence of compression artifacts in the decoded sequences, which, if not properly addressed, can degrade transcoding efficiency by causing suboptimal bit allocation and disrupting core coding processes. In this article, a learned video transcoding framework (LVT) is proposed to optimize video transcoding, leveraging coding priors from the input bitstream to guide the transcoding process. In the framework, to mitigate the adverse effects of compression artifacts, a Coding Priors-Guided Spatial Feature Transform module is designed, which utilizes coding prior features to adaptively modulate intermediate features through spatial affine transformations, enhancing bit allocation and suppressing artifacts. Additionally, a Coding Priors-Guided Quality Adapter module is proposed to generate a compression degradation representation using coding priors, which dynamically interacts with intermediate features to enable the network to perceive and adapt to different levels of degradation in the input video. Furthermore, a Motion Vectors-Guided Flow Refinement module is proposed to reduce prediction errors caused by artifacts. It refines optical flow predictions by using motion vectors from the bitstream as auxiliary information. Extensive experiments demonstrate that our framework outperforms both existing traditional and learned video codecs in transcoding performance, achieving an average bitrate saving of 20.3% compared to the H.266/VVC reference software VTM under the practical YUV420 setting measured with PSNR.
Nianxiang Fu, Daiqin Yang, Zhenan Lin et al.· ACM Trans. Multim. Comput. C...· 0 citations
This study evaluates the viability of deep learningbased super-resolution (SR) for enhancing real-time video streaming. We compare a traditional HEVC-compressed streaming pipeline against an AI-assisted framework that transmits lowresolution video and reconstructs high-resolution output at the client using the Efficient Sub-Pixel Convolutional Neural Network (ESPCN) algorithm. The evaluation spans multiple resolutions (1080p to 4K), motion types, and processing methodologies (subprocess vs. frame-by-frame). Experimental results indicate that while AI-assisted streaming significantly reduces transmission bandwidth, it introduces substantial computational overhead, with completion times increasing by over 300% compared to traditional methods. Furthermore, quality assessments using PSNR and VMAF metrics reveal that most AI-assisted streams failed to meet the target quality range (30–50 dB PSNR), with high-resolution 4K streams experiencing up to 24.58% frame loss during real-time inference. These findings highlight critical performance barriers in utilizing current SR algorithms for standard consumer-grade real-time video delivery.
Emma Hubbell, Hao Wu, Yulei Pang· 2026 International Conferenc...· 0 citations
Scalable video coding (SVC) encodes a video into a layered bitstream consisting of a base layer and one or multiple enhancement layers, enabling decoding at different bitrate/quality/resolution operating points to accommodate diverse device capabilities and network conditions. Due to its practical flexibility, SVC has been incorporated into major video coding standards and has recently attracted growing interest for both scene-agnostic and scene-adaptive neural video codecs. Among the latter, Implicit neural representation (INR) based codecs achieve compression by overfitting a compact neural network to an individual video, offering fast decoding and competitive coding efficiency compared to scene-agnostic neural codecs. However, research on scalable INR-based compression remains in its infancy: these methods support scalable coding by introducing additional network layers, which couple the bitrate with the decoding complexity and also cannot achieve comparable performance with strong scalable/non-scalable codecs. In this context, this paper proposes S-NVRC, a scalable INR-based video codec that jointly supports fine-grained bitrate and decoding complexity scalability from a single embedded bitstream. It adopts a coarse-to-fine prefix for feature grids and a nested prefix for network layers, which scale bitrate and decoding complexity, respectively. The proposed S-NVRC spans a wide range of bitrate and decoding-complexity using a single encoding (training) and outperforms SHM 12.4 and the multi-layer VTM-20.0, by 43.7% and 5.6% in BD-rate on the UVG dataset, while also providing flexible complexity scalability. Implemented code will be provided.
A frequency-aware compressed video quality enhancement framework that improves visual quality by adaptively enhancing high-frequency details and texture structures and reduces compression artifacts and enhances perceptual detail quality compared to existing approaches is proposed.
This paper proposes BinRVR, a binarized RAW video restoration framework that reduces computation and parameters by approximately 96% while incurring only about 4% performance degradation, and develops a Distribution-Aware Binarized Convolution (DAB-Conv) that leverages the statistics of full-precision activations to mitigate quantization errors.
Tianyu Zhu, Ying Fu, Hesong Li et al.· IEEE Transactions on Pattern...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.