Wei Jiang 1 Junrui Li 2 , Jiwei Xu 2 , 1 National Police University for Criminal Justice, Baoding, China, 2Beijing University of Chemical Technology, Beijing, China.
Achieving high-efficiency video compression while preserving perceptual quality is a persistent challenge, especially with the growing demand for high-resolution and high-frame-rate content. In this paper, we propose a novel compression framework that focuses on region-of-interest awareness and per-ceptual optimization to enhance both compression efficiency and visual fidelity. The core of our method lies in identifying foreground regions—treated as regions of interest—using an object detection algorithm prior to encoding. These regions receive prioritized treatment during compression, with finer quantization control guided by a deep learning model. This model, based on a convolutional neural network, dynamically predicts quantization levels for each coding block by incorporating both spatial features and perceptual quality metrics. To further improve encoding performance, we reformulate traditional rate-distortion op-timization by introducing data-driven models that relate bitrate and visual quality to quantization levels. These models serve to generate high-quality training labels and guide quantization decisions during in-ference. Additionally, region-aware encoding control adapts quantization granularity based on the size and significance of detected objects. Experimental results demonstrate that the proposed approach signif-icantly reduces bitrate—achieving an maximum saving of around 19%—while maintaining stable and high perceptual video quality, outperforming conventional video coding techniques and recent learning-based methods.
Copyright © AIFZ 2026