Token-Selective Vision Transformer for fine-grained image recognition of marine organisms-Reference-Cited by-同舟云学术

Token-Selective Vision Transformer for fine-grained image recognition of marine organisms

Published:2023-04-25 Issue: Volume:10 Page:
ISSN:2296-7745
Container-title:Frontiers in Marine Science
language:
Short-container-title:Front. Mar. Sci.

Author:

Si Guangzhe,Xiao Ying,Wei Bin,Bullock Leon Bevan,Wang Yueyue,Wang Xiaodong

Abstract

IntroductionThe objective of fine-grained image classification on marine organisms is to distinguish the subtle variations in the organisms so as to accurately classify them into subcategories. The key to accurate classification is to locate the distinguishing feature regions, such as the fish’s eye, fins, or tail, etc. Images of marine organisms are hard to work with as they are often taken from multiple angles and contain different scenes, additionally they usually have complex backgrounds and often contain human or other distractions, all of which makes it difficult to focus on the marine organism itself and identify its most distinctive features.Related workMost existing fine-grained image classification methods based on Convolutional Neural Networks (CNN) cannot accurately enough locate the distinguishing feature regions, and the identified regions also contain a large amount of background data. Vision Transformer (ViT) has strong global information capturing abilities and gives strong performances in traditional classification tasks. The core of ViT, is a Multi-Head Self-Attention mechanism (MSA) which first establishes a connection between different patch tokens in a pair of images, then combines all the information of the tokens for classification.MethodsHowever, not all tokens are conducive to fine-grained classification, many of them contain extraneous data (noise). We hope to eliminate the influence of interfering tokens such as background data on the identification of marine organisms, and then gradually narrow down the local feature area to accurately determine the distinctive features. To this end, this paper put forwards a novel Transformer-based framework, namely Token-Selective Vision Transformer (TSVT), in which the Token-Selective Self-Attention (TSSA) is proposed to select the discriminating important tokens for attention computation which helps limits the attention to more precise local regions. TSSA is applied to different layers, and the number of selected tokens in each layer decreases on the basis of the previous layer, this method gradually locates the distinguishing regions in a hierarchical manner.ResultsThe effectiveness of TSVT is verified on three marine organism datasets and it is demonstrated that TSVT can achieve the state-of-the-art performance.

Publisher

Frontiers Media SA

Subject

Ocean Engineering,Water Science and Technology,Aquatic Science,Global and Planetary Change,Oceanography

Reference57 articles.

1. Fish recognition based on robust features extraction from size and shape measurements using neural network;Alsmadi;Comput. Sci.,2010

2. Fish classification based on robust features extraction from color signature using back-propagation classifier;Alsmadi;Comput. Sci.,2011

3. Bird species categorization using pose normalized deep convolutional nets. in;Branson;Br. Mach. Vision Conference.,2014

4. End-to-end object detection with transformers. in;Carion;Eur. Conf. Comput. Vision.,2020

5. The devil is in the channels: mutual-channel loss for fine-grained image classification;Chang;IEEE Trans. Image Process.,2020

Cited by 5 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Underwater species classification using deep learning technique;Revista Română de Informatică și Automatică;2024-06-21

2. Survey of automatic plankton image recognition: challenges, existing solutions and future perspectives;Artificial Intelligence Review;2024-04-12

3. A dual-branch feature fusion neural network for fish image fine-grained recognition;The Visual Computer;2024-04-10

4. Duet of ViT and CNN: multi-scale dual-branch network for fine-grained image classification of marine organisms;Intelligent Marine Technology and Systems;2024-01-30

5. Enhancing Organizing Pneumonia Diagnosis: A Novel Super-token Transformer Approach for Masson Body Segmentation;In Vivo;2024