欢迎访问中国科学院大学学报,今天是

中国科学院大学学报 ›› 2026, Vol. 43 ›› Issue (4): 433-443.DOI: 10.7523/j.ucas.2024.042

• 数学与物理学 •    下一篇

基于伽马散度的半监督自训练图像分类

毛紫茜, 张三国()   

  1. 中国科学院大学数学科学学院,北京 100049
    中国科学院大数据挖掘与知识管理重点实验室,北京 100049
  • 收稿日期:2023-10-18 修回日期:2024-04-22 发布日期:2024-06-04
  • 通讯作者: 张三国
  • 基金资助:
    国家自然科学基金(12171454);国家自然科学基金(U19B2940);中央高校基本科研业务费专项资助

Semi-supervised self-training image classification method based on gamma divergence

Ziqian MAO, Sanguo ZHANG()   

  1. School of Mathematical Sciences,University of Chinese Academy of Sciences,Beijing 100049,China
    Key Laboratory of Big Data Mining and Knowledge Management,Chinese Academy of Sciences,Beijing 100049,China
  • Received:2023-10-18 Revised:2024-04-22 Published:2024-06-04
  • Contact: Sanguo ZHANG

摘要:

近年来,各种半监督自训练方法在解决图像分类问题时取得了巨大进展,受到人们广泛关注。然而,即使是主流前沿的FixMatch方法,在模型训练过程中也会出现伪标签错误累积导致的训练不稳定和类别间极端性能失衡问题。为解决该问题,在现有FixMatch方法的基础上引入具有稳健性质的伽马散度,并提出基于伽马散度的半监督自训练方法。该方法将伽马散度以正则项的形式添加到训练模型的损失函数中,对错误记忆的伪标签进行纠偏处理。在人工添加污染的模拟数据集和CIFAR-10数据集上进行不同方法的对比实验,验证了GammaFixMatch方法处理伪标签错误累积问题的优越性。

关键词: 图像分类, 半监督, 自训练, 伽马散度, 深度学习, FixMatch

Abstract:

In recent years, various semi-supervised self-training methods have made significant progress in solving image classification problems and have attracted widespread attention. However, even the state-of-the-art methods like FixMatch can suffer from training instability and extreme performance imbalance between classes due to the accumulation of pseudo-labeling errors. To address this issue, this paper introduces the concept of robust gamma divergence on top of the FixMatch method and proposes a semi-supervised self-training approach based on robust gamma divergence. In this method, we incorporate gamma divergence as a regularization term to the loss function of the training model to correct the mislabeled instances caused by error accumulation. Additionally, we conduct comparative experiments on artificially polluted simulated datasets and the CIFAR-10 dataset to demonstrate the superiority of the GammaFixMatch method in handling the problem of pseudo-label error accumulation.

Key words: image classification, semi-supervised, self-training, gamma divergence, deep learning, FixMatch

中图分类号: