[1] Zhang Chenglong, Fang Yang, Liang Xinyan, et al. Efficient multi-view unsupervised feature selection with adaptive structure learning and inference[C]//Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (IJCAI-24). 2024: 5443-5452.[1] Zhang Chenglong, Fang Yang, Liang Xinyan, et al. Efficient multi-view unsupervised feature selection with adaptive structure learning and inference[C]//Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence (IJCAI-24). 2024: 5443-5452.
[2] Caesar H, Bankiti V, Lang A H, et al. nuscenes: A multimodal dataset for autonomous driving[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2020: 11621-11631.
[3] Zou Xin, Tang Chang, Zheng Xiao, et al. Dpnet: Dynamic poly-attention network for trustworthy multi-modal classification[C]//Proceedings of the 31st ACM international conference on multimedia. 2023: 3550-3559.
[4] Zhou Hai, Xue Zhe, Liu Ying, et al. Rtmc: a rubost trusted multi-view classification framework[C]//2023 IEEE international conference on multimedia and expo (ICME). IEEE, 2023: 576-581.
[5] Zhou Hai, Xue Zhe, Liu Ying, et al. Calm: An enhanced encoding and confidence evaluating framework for trustworthy multi-view learning[C]//Proceedings of the 31st ACM International Conference on Multimedia. 2023: 3108-3116.
[6] Xie Mengyao, Han Zongbo, Zhang Changqing, et al. Exploring and exploiting uncertainty for incomplete multi-view classification[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023: 19873-19882.
[7] Ma Qian, Gu Yu, Zhang Tiancheng, et al. A heterogeneous multi-source multi-mode sensory data acquisition method based on data quality[J]. Chinese Journal of Computers, 2013, 36(10): 2120-2131.
[8] Peterson L E. K-nearest neighbor[J]. Scholarpedia, 2009, 4(2): 1883.
[9] Peng Chong, Kang Kehan, Chen Yongyong, et al. Fine-grained essential tensor learning for robust multi-view spectral clustering[J]. IEEE Transactions on Image Processing, 2024.
[10] Bian Jintang, Xie Xiaohua, Lai Jianhuang, et al. Multi-view contrastive clustering via integrating graph aggregation and confidence enhancement[J]. Information Fusion, 2024, 108: 102393.
[11] 周嘉文, 郑小盈, 祝永新, 林思敏, 陈凌曜, 曾洪斌, 郭俞, 王馨莹. 多头自注意力与双线性池化融合的心肌缺血影像分类[J]. 计算机工程, ZHOU Jiawen, ZHENG Xiaoying, ZHU Yongxin, LIN Simin, CHEN Lingyao, ZENG Hongbin, GUO Yu, WANG Xinying. Myocardial Ischemia Image Classification via Fusion of Multi-Head Self-Attention and Bilinear Pooling[J]. Computer Engineering, 2025, 51(11): 246-257.
[12] Zhou Lihua, Du Guowang, Lü Kevin, et al. A survey and an empirical evaluation of multi-view clustering approaches[J]. ACM Computing Surveys, 2024, 56(7): 1-38.
[13] Zheng Xiao, Wang Minhui, Huang Kai, et al. Global and cross-modal feature aggregation for multi-omics data classification and application on drug response prediction[J]. Information Fusion, 2024, 102: 102077.
[14] Kiela D, Grave E, Joulin A, et al. Efficient large-scale multi-modal classification[C]//Proceedings of the AAAI conference on artificial intelligence. 2018, 32(1).
[15] Panda R, Chen C F R, Fan Q, et al. Adamml: Adaptive multi-modal learning for efficient video recognition[C]//Proceedings of the IEEE/CVF international conference on computer vision. 2021: 7576-7585.
[16] Wang Hao, Yang Yan, Liu Bing. GMC: Graph-based multi-view clustering[J]. IEEE Transactions on Knowledge and Data Engineering, 2019, 32(6): 1116-1129.
[17] Abdar M, Pourpanah F, Hussain S, et al. A review of uncertainty quantification in deep learning: Techniques, applications and challenges[J]. Information fusion, 2021, 76: 243-297.
[18] 周世向, 于凯. 基于密集协同注意力的多模态情感分析[J]. 计算机工程, 2025, 51(11): 144-151. ZHOU Shixiang, YU Kai. Multimodal Sentiment Analysis Based on Dense Co-Attention[J]. Computer Engineering, 2025, 51(11): 144-151.
[19] Wang Chunwei, Ma Chao, Zhu Ming, et al. Pointaugmenting: Cross-modal augmentation for 3d object detection[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2021: 11794-11803.
[20] Duro-Castano A, Borrás C, Herranz-Pérez V, et al. Targeting Alzheimer's disease with multimodal polypeptide-based nanoconjugates[J]. Science Advances, 2021, 7(13): eabf9180.
[21] Siejka-Zielińska P, Cheng J, Jackson F, et al. Cell-free DNA TAPS provides multimodal information for early cancer detection[J]. Science advances, 2021, 7(36): eabh0534.
[22] Antorán J, Allingham J, Hernández-Lobato J M. Depth uncertainty in neural networks[J]. Advances in neural information processing systems, 2020, 33: 10620-10634.
[23] Havasi M, Jenatton R, Fort S, et al. Training independent subnetworks for robust prediction[J]. arXiv preprint arXiv:2010.06610, 2020.
[24] Han Zongbo, Yang Fan, Huang Junzhou, et al. Multimodal dynamics: Dynamical fusion for trustworthy multimodal classification[C]//Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2022: 20707-20717.
[25] Geng Yu, Han Zongbo, Zhang Changqing, et al. Uncertainty-aware multi-view representation learning[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2021, 35(9): 7545-7553.
[26] Poria S, Cambria E, Gelbukh A. Deep convolutional neural network textual features and multiple kernel learning for utterance-level multimodal sentiment analysis[C]//Proceedings of the 2015 conference on empirical methods in natural language processing. 2015: 2539-2544.
[27] Huang Yu, Du Chenzhuang, Xue Zihui, et al. What makes multi-modal learning better than single (provably)[J]. Advances in Neural Information Processing Systems, 2021, 34: 10944-10956.
[28] Kiela D, Bhooshan S, Firooz H, et al. Supervised multimodal bitransformers for classifying images and text[J]. arXiv preprint arXiv:1909.02950, 2019.
[29] Han Zongbo, Zhang Changqing, Fu Huazhu, et al. Trusted multi-view classification with dynamic evidential fusion[J]. IEEE transactions on pattern analysis and machine intelligence, 2022, 45(2): 2551-2566.
[30] Natarajan P, Wu S, Vitaladevuni S, et al. Multimodal feature fusion for robust event detection in web videos[C]//2012 IEEE conference on computer vision and pattern recognition. IEEE, 2012: 1298-1305.
[31] Simonyan K, Zisserman A. Two-stream convolutional networks for action recognition in videos[J]. Advances in neural information processing systems, 2014, 27.
[32] Wu Guanyao, Liu Haoyu, Fu Hongming, et al. Every SAM Drop Counts: Embracing Semantic Priors for Multi-Modality Image Fusion and Beyond[C]//Proceedings of the Computer Vision and Pattern Recognition Conference. 2025: 17882-17891.
[33] Gao Jingsheng, Ruan Jiacheng, Xiang Suncheng, et al. Lamm: Label alignment for multi-modal prompt learning[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2024, 38(3): 1815-1823.
[34] Burchi M, Timofte R. Audio-visual efficient conformer for robust speech recognition[C]//Proceedings of the IEEE/CVF winter conference on applications of computer vision. 2023: 2258-2267.
[35] Tsai Y H H, Bai Shaojie, Liang P P, et al. Multimodal transformer for unaligned multimodal language sequences[C]//Proceedings of the conference. Association for computational linguistics. Meeting. 2019, 2019:6558.
[36] Wang Yikai, Huang Wenbing, Sun Fuchun, et al. Deep multimodal fusion by channel exchanging[J]. Advances in neural information processing systems, 2020, 33: 4835-4845.
[37] 孙圆, 王康平, 赵鸣博. 基于多提示和图文对比学习的服装检索[J]. 计算机工程, SUN Yuan, WANG Kangping, ZHAO Mingbo. Clothing Retrieval Based on Multiple Prompts and Contrastive Image-Text Learning[J]. Computer Engineering, 2026, 52(2): 322-330.
[38] 郭浩, 李欣奕, 唐九阳, 等.自适应特征融合的多模态实体对齐研究. 自动化学报, 2024, 50(4): 758?770.
(Guo Hao, Li Xinyi, Tang Jiuyang, et al. Adaptive feature fusion for multi-modal entity alignment. Acta Automatica Sinica, 2024, 50(4): 758?770.)
[39] 马茜, 谷峪, 张天成, 等.一种基于数据质量的异构多源多模态感知数据获取方法. 计算机学报, 2013, 36(10): 2120-2131.
(Ma Qian, GU Yu, Zhang Tiancheng, et al. A heterogeneous multi-source multi-mode sensory data acquisition method based on data quality. Chinese Journal of Computers, 2013, 36(10): 2120-2131.)
[40] Wah C, Branson S, Welinder P, et al. The caltech-ucsd birds-200-2011 dataset[C]. California institute of technology, 2011
[41] Wang Weiran, Arora R, Livescu K, et al. On deep multi-view representation learning[C]//International conference on machine learning. PMLR, 2015: 1083-1092.
[42] Fei-Fei L, Fergus R, Perona P. Learning generative visual models from few training examples: An incremental bayesian approach tested on 101 object categories[C]//2004 conference on computer vision and pattern recognition workshop. IEEE, 2004: 178-178.
[43] Gross R, Matthews I, Cohn J, et al. Multi-pie[J]. Image and vision computing, 2010, 28(5): 807-813
[44] Lampert C H, Nickisch H, Harmeling S. Attribute-based classification for zero-shot visual object categorization[J]. IEEE transactions on pattern analysis and machine intelligence, 2013, 36(3): 453-465.
[45] Andrew G, Arora R, Bilmes J, et al. Deep canonical correlation analysis[C]//International conference on machine learning. PMLR, 2013: 1247-1255.
[46] Hwang H J, Kim G H, Hong S, et al. Multi-view representation learning via total correlation objective[J]. Advances in Neural Information Processing Systems, 2021, 34: 12194-12207.
[47] Gal Y, Ghahramani Z. Bayesian convolutional neural networks with Bernoulli approximate variational inference[J]. arXiv preprint arXiv:1506.02158, 2015.
[48] Lakshminarayanan B, Pritzel A, Blundell C. Simple and scalable predictive uncertainty estimation using deep ensembles[J]. Advances in neural information processing systems, 2017, 30.
[49] Sensoy M, Kaplan L, Kandemir M. Evidential deep learning to quantify classification uncertainty[J]. Advances in neural information processing systems, 2018, 31.
[50] Liu Wei, Chen Yufei, Yue Xiaodong, et al. Safe multi-view deep classification[C]//Proceedings of the AAAI Conference on Artificial Intelligence. 2023, 37(7): 8870-8878.
[51] Yue Xiaodong, Dong Zhicheng, Chen Yufei, et al. Evidential dissonance measure in robust multi-view classification to resist adversarial attack[J]. Information Fusion, 2025, 113: 102605.
[52] Li C, Mo L, Kwoh C K, et al. Noise-robust multi-view graph neural network for fault diagnosis of rotating machinery[J]. Mechanical Systems and Signal Processing, 2025, 224: 112025.
[53] Liang V W, Zhang Y, Kwon Y, et al. Mind the gap: Understanding the modality gap in multi-modal contrastive representation learning[J]. Advances in Neural Information Processing Systems, 2022, 35: 17612-17625.
[54] Ge S, Chen Y. Confidence-aware multimodal learning for trustworthy fake news detection[J]. INFORMS Journal on Computing, 2025.
|