作者投稿和查稿 主编审稿 专家审稿 编委审稿 远程编辑

计算机工程 ›› 2026, Vol. 52 ›› Issue (10): 227-241. doi: 10.19678/j.issn.1000-3428.0252123

• 计算机视觉与图形图像处理 • 上一篇    

仿双目竞争的无参考立体图像质量评价

徐少平, 王子超, 唐祎玲, 熊思龙   

  1. 南昌大学数学与计算机学院, 江西 南昌 330031
  • 收稿日期:2025-02-12 修回日期:2025-04-10 发布日期:2026-09-29
  • 作者简介:徐少平,男,教授、博士,主研方向为数字图像处理;王子超,硕士研究生;唐祎玲(通信作者),副教授、博士,E-mail:tangyiling@ncu.edu.cn;熊思龙,硕士研究生。
  • 基金资助:
    国家自然科学基金(62162042,62162043)。

No-Reference Stereoscopic Image Quality Assessment Simulating Binocular Rivalry

XU Shaoping, WANG Zichao, TANG Yiling, XIONG Silong   

  1. School of Mathematics and Computer Sciences, Nanchang University, Nanchang 330031, Jiangxi, China
  • Received:2025-02-12 Revised:2025-04-10 Published:2026-09-29

摘要: 人类视觉皮层采用分层结构,其中双目融合与双目竞争首先发生在低级视觉区域,但当前基于深度学习的立体图像质量评价(SIQA)模型普遍采用在网络的不同层次上融合左右视点图像特征来估计立体图像质量值,对人类低级视觉区域感知的模拟程度存在不足。鉴于此,本文提出了一种仿双目竞争的立体图像质量评价方法。首先,模拟双目视觉竞争现象,构建了一个基于无监督方法的双目图像融合模型。通过左右视点图像的梯度幅值响应来评估图像降质程度,确定左右视点图像的融合权重,并利用深度卷积神经网络(CNN)对输入图像先验知识的获取能力,建立基于编码器-解码器架构的无监督图像生成网络,以左右视点两幅图像作为学习对象,实现左右视点图像的融合。其次,利用在大规模图像数据库上预训练的ResNet50模型从融合图像中提取质量感知特征,并构建了一个基于支持向量回归(SVR)的特征质量映射模型来估计立体图像的质量值。实验结果显示,在4个经典立体图像基准测试数据库上,所提方法在皮尔逊线性相关系数(PLCC)和斯皮尔曼等级相关系数(SROCC)两个评价指标上均超过了0.96,并且均方根误差(RMSE)均优于对比方法。这表明提出的基于无监督双目图像融合的方法能够有效模拟双目视觉效应,从而显著提高立体图像质量评价的准确性。

关键词: 立体图像质量评价, 双目图像融合, 双目竞争, 深度图像先验, 无监督学习

Abstract: The human visual cortex has a hierarchical structure in which binocular fusion and binocular rivalry first occur in low-level visual areas. However, current deep learning-based Stereoscopic Image Quality Assessment (SIQA) models generally estimate the quality values of stereoscopic images by fusing features of left- and right-view images at different levels of the network, resulting in insufficient simulation of the perceptual mechanisms in the low-level visual areas of the human visual cortex. To address this issue, this study proposes a stereoscopic image quality assessment method that simulates binocular rivalry. First, to simulate binocular visual rivalry, a binocular image fusion model based on an unsupervised approach is constructed by leveraging the ability of deep Convolutional Neural Networks (CNNs) to acquire prior knowledge of the input images. Considering the left and right views as learning targets, this model is built on an encoder—decoder architecture, and it simulates the binocular fusion process in the human visual system. The gradient magnitude responses of the left and right images are used to evaluate the degree of image degradation and to determine the fusion weights of the left and right views, which simulate the binocular rivalry phenomenon. Second, a ResNet50 model pretrained on a large-scale image database is used to extract quality-aware features from the fused image and construct a feature-quality mapping model based on Support Vector Regression (SVR) for estimating the quality value of the stereoscopic image. Experimental results on four classic stereoscopic image benchmark databases show that the proposed method achieves over 0.96 in both the Pearson Linear Correlation Coefficient (PLCC) and Spearman Rank Order Correlation Coefficient (SROCC), and its Root Mean Square Error (RMSE) is superior to that of the compared methods. These results indicate that the proposed unsupervised binocular image fusion method can effectively mimic binocular visual effects, thus significantly improving the accuracy of stereoscopic image quality assessment.

Key words: Stereoscopic Image Quality Assessment (SIQA), binocular image fusion, binocular rivalry, Deep Image Prior (DIP), unsupervised learning

中图分类号: