ConViT: improving vision transformers with soft convolutional inductive biases*-Reference-Cited by-同舟云学术

ConViT: improving vision transformers with soft convolutional inductive biases*

Published:2022-11-01 Issue:11 Volume:2022 Page:114005
ISSN:1742-5468
Container-title:Journal of Statistical Mechanics: Theory and Experiment
language:
Short-container-title:J. Stat. Mech.

Author:

d’Ascoli Stéphane,Touvron Hugo,Leavitt Matthew L,Morcos Ari S,Biroli Giulio,Sagun Levent

Abstract

Abstract Convolutional architectures have proven to be extremely successful for vision tasks. Their hard inductive biases enable sample-efficient learning, but come at the cost of a potentially lower performance ceiling. Vision transformers rely on more flexible self-attention layers, and have recently outperformed CNNs for image classification. However, they require costly pre-training on large external datasets or distillation from pre-trained convolutional networks. In this paper, we ask the following question: is it possible to combine the strengths of these two architectures while avoiding their respective limitations? To this end, we introduce gated positional self-attention (GPSA), a form of positional self-attention which can be equipped with a ‘soft’ convolutional inductive bias. We initialize the GPSA layers to mimic the locality of convolutional layers, then give each attention head the freedom to escape locality by adjusting a gating parameter regulating the attention paid to position versus content information. The resulting convolutional-like ViT architecture, ConViT, outperforms the DeiT (Touvron et al 2020 arXiv:2012.12877) on ImageNet, while offering a much improved sample efficiency. We further investigate the role of locality in learning by first quantifying how it is encouraged in vanilla self-attention layers, then analyzing how it has escaped in GPSA layers. We conclude by presenting various ablations to better understand the success of the ConViT. Our code and models are released publicly at https://github.com/facebookresearch/convit.

Publisher

IOP Publishing

Subject

Statistics, Probability and Uncertainty,Statistics and Probability,Statistical and Nonlinear Physics

Link

https://iopscience.iop.org/article/10.1088/1742-5468/ac9830/pdf

Reference48 articles.

1. Transferring inductive biases through knowledge distillation;Abnar,2020

2. Homotopy analysis for tensor PCA;Anandkumar,2016

3. Neural machine translation by jointly learning to align and translate;Bahdanau,2014

4. Attention augmented convolutional networks;Bello,2019

5. End-to-end object detection with transformers;Carion,2020

Cited by 389 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. ReViT: Enhancing vision transformers feature diversity with attention residual connections;Pattern Recognition;2024-12

2. Back-to-Bones: Rediscovering the role of backbones in domain generalization;Pattern Recognition;2024-12

3. Effect of spatial scale, color infrared and sample size on learning poverty from aerial images;Remote Sensing Applications: Society and Environment;2024-11

4. Deblurring masked image modeling for ultrasound image analysis;Medical Image Analysis;2024-10

5. Vision transformer promotes cancer diagnosis: A comprehensive review;Expert Systems with Applications;2024-10