Image–text coherence and its implications for multimodal AI-Reference-Cited by-同舟云学术

Image–text coherence and its implications for multimodal AI

Published:2023-05-15 Issue: Volume:6 Page:
ISSN:2624-8212
Container-title:Frontiers in Artificial Intelligence
language:
Short-container-title:Front. Artif. Intell.

Author:

Alikhani Malihe,Khalid Baber,Stone Matthew

Abstract

Human communication often combines imagery and text into integrated presentations, especially online. In this paper, we show how image–text coherence relations can be used to model the pragmatics of image–text presentations in AI systems. In contrast to alternative frameworks that characterize image–text presentations in terms of the priority, relevance, or overlap of information across modalities, coherence theory postulates that each unit of a discourse stands in specific pragmatic relations to other parts of the discourse, with each relation involving its own information goals and inferential connections. Text accompanying an image may, for example, characterize what's visible in the image, explain how the image was obtained, offer the author's appraisal of or reaction to the depicted situation, and so forth. The advantage of coherence theory is that it provides a simple, robust, and effective abstraction of communicative goals for practical applications. To argue this, we review case studies describing coherence in image–text data sets, predicting coherence from few-shot annotations, and coherence models of image–text tasks such as caption generation and caption evaluation.

Publisher

Frontiers Media SA

Subject

Artificial Intelligence

Reference57 articles.

1. “Applying discourse semantics and pragmatics to co-reference in picture sequences,”;Abusch,2013

2. “Cite: a corpus of image-text discourse relations,”;Alikhani;Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers),2019

3. “Cross-modal coherence for text-to-image retrieval,”;Alikhani;Proceedings of the 36th AAAI Conference on Artificial Intelligence, Vol. 10,2022

4. “Cross-modal coherence modeling for caption generation,”;Alikhani,2020

5. “Exploring coherence in visual explanations,”;Alikhani,2018

Cited by 2 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. How verbal text guides the interpretation of advertisement images: a predictive typology of verbal anchoring;Communication Theory;2024-07-22

2. Large Language Models: A Historical and Sociocultural Perspective;Cognitive Science;2024-03