Author Login Chief Editor Login Reviewer Login Editor Login Remote Office

Computer Engineering ›› 2026, Vol. 52 ›› Issue (9): 81-91. doi: 10.19678/j.issn.1000-3428.0252092

• Computational Intelligence and Pattern Recognition • Previous Articles     Next Articles

Tabular Ordinal Classification Based on Ordinal Entropy Optimization Model

LUO Zhengdong1,2,3, ZHANG Guohao1,2,3, HAN Yunfei1,2,3, WANG Yi1,2,3, ZHOU Xi1,2,3,*()   

  1. 1. Xinjiang Technical Institute of Physics and Chemistry, Chinese Academy of Sciences, Urumqi 830011, Xinjiang, China
    2. University of Chinese Academy of Sciences, Beijing 100049, China
    3. Xinjiang Key Laboratory of Minority Speech and Language Information Processing, Chinese Academy of Sciences, Urumqi 830011, Xinjiang, China
  • Received:2025-01-24 Revised:2025-03-14 Online:2026-09-15 Published:2025-05-09
  • Contact: ZHOU Xi

基于有序熵优化模型的表格有序分类

罗正东1,2,3, 张国昊1,2,3, 韩云飞1,2,3, 王轶1,2,3, 周喜1,2,3,*()   

  1. 1. 中国科学院新疆理化技术研究所, 新疆 乌鲁木齐 830011
    2. 中国科学院大学, 北京 100049
    3. 中国科学院新疆民族语音语言信息处理重点实验室, 新疆 乌鲁木齐 830011
  • 通讯作者: 周喜
  • 作者简介:

    罗正东, 男, 博士研究生, 主研方向为数据挖掘、表格有序分类

    张国昊, 博士研究生

    韩云飞(CCF会员), 副研究员、博士

    王轶(CCF高级会员), 研究员、博士

    周喜(通信作者), 研究员、博士

  • 基金资助:
    新疆维吾尔自治区重大科技专项(2023A01006); 新疆维吾尔自治区重点研发计划(2023B01028); 新疆维吾尔自治区"天山英才"项目(2022TSYCLJ0035); 新疆维吾尔自治区"天山英才"项目(2023TSYCTD0011); 新疆维吾尔自治区"天山英才"项目(2023TSYCCX0046); 中国科学院青年创新促进会项目(2021434)

Abstract:

The existing methods for tabular data prediction primarily focus on classical classification and regression tasks. However, a type of data in the tabular data domain contains labels with an ordinal relationship, and its prediction task is called tabular ordinal classification. Current methods for tabular ordinal classification rely primarily on retrieving similar features and augmenting sample feature representations by fusing features similar to the ordinal distance between classes. However, the existing methods neglect the full utilization of label ordinal knowledge. To address this, a method based on ordinal label entropy optimization is proposed, which effectively guides the model in learning ordinal information by mining the ordinal entropy embedded in the label order knowledge. First, an ordinal entropy calculation module is established that quantifies ordinal entropy based on the rank differences between the predicted and true labels. Through step-by-step analysis and derivation, the ordinal label entropy is designed as a novel rank loss function, which is introduced as a regularization term in the model. This encourages the model to learn the ordinal relationship between labels and reduces information loss caused by unordered predictions. The ordinal entropy optimized ranking loss function is combined with the original loss function of the model to improve its predictive ability. Finally, experiments on multiple ordinal tabular datasets reveal that this method outperforms various baseline models, demonstrating the effectiveness and advantages of the ordinal entropy optimization model in tabular ordinal classification tasks.

Key words: tabular data, ordinal classification, ordinal information, ordinal entropy, rank loss

摘要:

现有表格数据预测方法主要聚焦于传统分类和回归的研究, 然而在表格数据领域中存在一种标签具有有序关系的数据类型, 其预测任务被称为表格有序分类。目前的表格有序分类方法主要采用检索相似特征的方式, 通过相似特征与类间有序距离融合增强样本特征表示。但现有方法忽略了标签有序知识的充分利用, 因此提出一种基于标签有序熵优化的方法, 通过挖掘标签有序知识中蕴含的有序熵, 有效指导模型学习有序信息。具体而言, 首先建立有序熵计算模块, 利用预测标签与真实标签之间的等级顺序差异量化有序熵。通过逐步分析和推导, 将标签有序熵设计为一种新颖的排序损失函数作为正则项引入模型, 鼓励模型学习标签等级顺序关系, 以减少无序预测带来的信息损失。然后, 将有序熵优化排序损失函数与模型原有损失函数相结合, 共同提升模型的预测能力。在多个有序表格数据集上的实验结果显示, 该方法相较于多种基线模型取得了性能提升, 充分证明了有序熵优化模型在表格有序分类任务中的有效性与优势。

关键词: 表格数据, 有序分类, 有序信息, 有序熵, 排序损失