Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-Training-Reference-Cited by-同舟云学术

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-Training

Published:2020-04-03 Issue:07 Volume:34 Page:11336-11344
ISSN:2374-3468
Container-title:Proceedings of the AAAI Conference on Artificial Intelligence
language:
Short-container-title:AAAI

Author:

Li Gen,Duan Nan,Fang Yuejian,Gong Ming,Jiang Daxin

Abstract

We propose Unicoder-VL, a universal encoder that aims to learn joint representations of vision and language in a pre-training manner. Borrow ideas from cross-lingual pre-trained models, such as XLM (Lample and Conneau 2019) and Unicoder (Huang et al. 2019), both visual and linguistic contents are fed into a multi-layer Transformer (Vaswani et al. 2017) for the cross-modal pre-training, where three pre-trained tasks are employed, including Masked Language Modeling(MLM), Masked Object Classification(MOC) and Visual-linguistic Matching(VLM). The first two tasks learn context-aware representations for input tokens based on linguistic and visual contents jointly. The last task tries to predict whether an image and a text describe each other. After pretraining on large-scale image-caption pairs, we transfer Unicoder-VL to caption-based image-text retrieval and visual commonsense reasoning, with just one additional output layer. We achieve state-of-the-art or comparable results on both two tasks and show the powerful ability of the cross-modal pre-training.

Publisher

Association for the Advancement of Artificial Intelligence (AAAI)

Subject

General Medicine

Cited by 311 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Cross-modal collaborative feature representation via Transformer-based multimodal mixers for RGB-T crowd counting;Expert Systems with Applications;2024-12

2. An end-to-end image-text matching approach considering semantic uncertainty;Neurocomputing;2024-11

3. Understanding mobile GUI: From pixel-words to screen-sentences;Neurocomputing;2024-10

4. BagFormer: Better cross-modal retrieval via bag-wise interaction;Engineering Applications of Artificial Intelligence;2024-10

5. The Implementation of Multimodal Large Language Models for Hydrological Applications: A Comparative Study of GPT-4 Vision, Gemini, LLaVa, and Multimodal-GPT;Hydrology;2024-09-11