A Survey of Vision-Language Pre-Trained Models-Reference-Cited by-同舟云学术

A Survey of Vision-Language Pre-Trained Models

Published:2022-07 Issue: Volume: Page:
ISSN:
Container-title:Proceedings of the Thirty-First International Joint Conference on Artificial Intelligence
language:
Short-container-title:

Author:

Du Yifan¹²,Liu Zikang¹,Li Junyi¹³,Zhao Wayne Xin¹²

Affiliation:

1. Renmin University of China

2. Beijing Key Laboratory of Big Data Management and Analysis Methods

3. University of Montreal

Abstract

As transformer evolves, pre-trained models have advanced at a breakneck pace in recent years. They have dominated the mainstream techniques in natural language processing (NLP) and computer vision (CV). How to adapt pre-training to the field of Vision-and-Language (V-L) learning and improve downstream task performance becomes a focus of multimodal learning. In this paper, we review the recent progress in Vision-Language Pre-Trained Models (VL-PTMs). As the core content, we first briefly introduce several ways to encode raw images and texts to single-modal embeddings before pre-training. Then, we dive into the mainstream architectures of VL-PTMs in modeling the interaction between text and image representations. We further present widely-used pre-training tasks, and then we introduce some common downstream tasks. We finally conclude this paper and present some promising research directions. Our survey aims to provide researchers with synthesis and pointer to related research.

Publisher

International Joint Conferences on Artificial Intelligence Organization

Cited by 35 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. SiamQuality: a ConvNet-based foundation model for photoplethysmography signals;Physiological Measurement;2024-08-01

2. Vision-Language Models for Vision Tasks: A Survey;IEEE Transactions on Pattern Analysis and Machine Intelligence;2024-08

3. Exploiting Pre-Trained Models and Low-Frequency Preference for Cost-Effective Transfer-based Attack;ACM Transactions on Knowledge Discovery from Data;2024-07-25

4. EASI-Tex: Edge-Aware Mesh Texturing from Single Image;ACM Transactions on Graphics;2024-07-19

5. Vesper: A Compact and Effective Pretrained Model for Speech Emotion Recognition;IEEE Transactions on Affective Computing;2024-07