
Multi‐label Image Classification via Coarse‐to‐Fine Attention*
Author(s) -
Lyu Fan,
Li Linyan,
Victor S. Sheng,
Fu Qiming,
Hu Fuyuan
Publication year - 2019
Publication title -
chinese journal of electronics
Language(s) - English
Resource type - Journals
SCImago Journal Rank - 0.267
H-Index - 25
eISSN - 2075-5597
pISSN - 1022-4653
DOI - 10.1049/cje.2019.07.015
Subject(s) - computer science , margin (machine learning) , artificial intelligence , pascal (unit) , image (mathematics) , pattern recognition (psychology) , deep neural networks , multi label classification , artificial neural network , machine learning , programming language
Great efforts have been made by using deep neural networks to recognize multi‐label images. Since multi‐label image classification is very complicated, many studies seek to use the attention mechanism as a kind of guidance. Conventional attention‐based methods always analyzed images directly and aggressively, which is difficult to well understand complicated scenes. We propose a global/local attention method that can recognize a multi‐label image from coarse to fine by mimicking how human‐beings observe images. Our global/local attention method first concentrates on the whole image, and then focuses on its local specific objects. We also propose a joint max‐margin objective function, which enforces that the minimum score of positive labels should be larger than the maximum score of negative labels horizontally and vertically. This function further improve our multi‐label image classification method. We evaluate the effectiveness of our method on two popular multi‐label image datasets ( i.e. , Pascal VOC and MS‐COCO). Our experimental results show that our method outperforms state‐of‐the‐art methods.