Multimodal Fine-Grained Transformer Model for Pest Recognition-Reference-Cited by-同舟云学术

Multimodal Fine-Grained Transformer Model for Pest Recognition

Published:2023-06-10 Issue:12 Volume:12 Page:2620
ISSN:2079-9292
Container-title:Electronics
language:en
Short-container-title:Electronics

Author:

Zhang Yinshuo¹²,Chen Lei¹,Yuan Yuan¹^ORCID

Affiliation:

1. Institute of Intelligent Machines, Hefei Institutes of Physical Science, Chinese Academy of Sciences, Hefei 230031, China

2. Science Island Branch, Graduate School of University of Science and Technology of China, Hefei 230026, China

Abstract

Deep learning has shown great potential in smart agriculture, especially in the field of pest recognition. However, existing methods require large datasets and do not exploit the semantic associations between multimodal data. To address these problems, this paper proposes a multimodal fine-grained transformer (MMFGT) model, a novel pest recognition method that improves three aspects of transformer architecture to meet the needs of few-shot pest recognition. On the one hand, the MMFGT uses self-supervised learning to extend the transformer structure to extract target features using contrastive learning to reduce the reliance on data volume. On the other hand, fine-grained recognition is integrated into the MMFGT to focus attention on finely differentiated areas of pest images to improve recognition accuracy. In addition, the MMFGT further improves the performance in pest recognition by using the joint multimodal information from the pest’s image and natural language description. Extensive experimental results demonstrate that the MMFGT obtains more competitive results compared to other excellent models, such as ResNet, ViT, SwinT, DINO, and EsViT, in pest recognition tasks, with recognition accuracy up to 98.12% and achieving 5.92% higher accuracy compared to the state-of-the-art DINO method for the baseline.

Funder

National Natural Science Foundation of China

National Basic Science Data Center

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Computer Networks and Communications,Hardware and Architecture,Signal Processing,Control and Systems Engineering

Link

https://www.mdpi.com/2079-9292/12/12/2620/pdf

Reference38 articles.

1. Review of deep learning: Concepts, CNN architectures, challenges, applications, future directions;Alzubaidi;J. Big Data,2021

2. Recognition pest by image-based transfer learning;Dawei;J. Sci. Food Agric.,2019

3. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., and Gelly, S. (2021, January 3–7). An image is worth 16 × 16 words: Transformers for image recognition at scale. Proceedings of the 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria.

4. Pest identification via deep residual learning in complex background;Cheng;Comput. Electron. Agric.,2017

5. Feature Reuse Residual Networks for Insect Pest Recognition;Ren;IEEE Access,2019

Cited by 3 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Overview of Pest Detection and Recognition Algorithms;Electronics;2024-07-30

2. Efficient agricultural pest classification using vision transformer with hybrid pooled multihead attention;Computers in Biology and Medicine;2024-07

3. Classification of Plant Leaf Disease Recognition Based on Self-Supervised Learning;Agronomy;2024-02-28