AI-enabled Industrial Innovation in Building Materials, Construction, Automotive and Transportation
SUN Yu, GAO Shu, ZHANG Yanxin
The video stream of maritime multifunctional navigation mark monitoring presents the distinctive scene characteristics of a dominant static sea surface background and sparse dynamic ship targets. Constrained by the low-bandwidth transmission of navigation mark-borne communication networks, compressed video streams experience severe quality degradation, including noise superposition, blocking artifacts, ringing artifacts, frame blurring, and inter-frame motion estimation bias. Existing general-purpose video quality enhancement methods are unable to jointly balance single-frame spatial detail recovery and multi-frame temporal information utilization, and therefore fall short of the requirements of the aforementioned scenario. To address this issue, a Content Adaptive Fusion-based Video Stream Quality Enhancement of Maritime Multifunctional Navigation Mark Monitoring (CAF-VSMM) method is proposed, in which a dual-branch fusion enhancement model is constructed. The single-frame enhancement branch, built on multiscale features and a spatio-temporal network, exploits the latent temporal information of a single-frame via virtual frames and integrates the pixel shuffle with a residual network to achieve multiscale spatial feature fusion, effectively suppressing compression artifacts and spatial degradation induced by the maritime environment. The multi-frame enhancement branch, based on a dynamic shifted window and inter-frame motion compensation, efficiently models long-range spatio-temporal dependencies and employs hierarchical offset estimation together with deformable convolution to achieve accurate inter-frame alignment in multiship scenes, thereby alleviating temporal blurring and motion misalignment. The Dynamic Feature Perception-based Content-Adaptive Fusion (DFP-CAF) module, based on dynamic feature perception, dynamically adjusts the feature weights of the two branches according to the video content, realizing an adaptive enhancement strategy that emphasizes detailed recovery for static backgrounds and temporal consistency for dynamic targets. Experimental results show that, on the self-built Port Cluster Waterway Monitoring Video Dataset (PC-WMVD), the Peak Signal to Noise Ratio (PSNR), Structural Similarity (SSIM), and Video Multimethod Evaluation Fusion (VMAF) of the proposed method reach 42.71 dB, 99.33%, and 77.14%, outperforming the second-best method by 3.90 dB, 1.19 percentage points, and 6.50 percentage points, respectively; on the public dataset JCT-VC, they reach 28.66 dB, 88.60%, and 55.40%, exceeding the second-best method by 1.63 dB, 5.42 percentage points, and 3.19 percentage points, respectively. The model has 37.52×106 parameters and 51.74×109 Floating-Point Operations Per Second (FLOPs) and runs at 41.52 frames per second, meeting the real-time requirements of waterway monitoring. The proposed method is tested in an intelligent monitoring system for the waterways of port clusters, where it significantly improves the identifiability of key targets, such as ships and navigation marks, and effectively supports downstream tasks, such as ship tracking, verifying its engineering practicability and application value.