Efficient Visual Metaphor Image Generation Based on Metaphor Understanding-Reference-Cited by-同舟云学术

Efficient Visual Metaphor Image Generation Based on Metaphor Understanding

Published:2024-04-16 Issue:3 Volume:56 Page:
ISSN:1573-773X
Container-title:Neural Processing Letters
language:en
Short-container-title:Neural Process Lett

Author:

Su Chang,Wang Xingyue,Liu Shupin,Chen Yijiang

Abstract

AbstractMetaphor has significant implications for revealing cognitive and thinking mechanisms. Visual metaphor image generation not only presents metaphorical connotations intuitively but also reflects AI’s understanding of metaphor through the generated images. This paper investigates the task of generating images based on text with visual metaphors. We explore metaphor image generation and create a dataset containing sentences with visual metaphors. Then, we propose a visual metaphor generation image framework based on metaphor understanding, which is more tailored to the essence of metaphor, better utilizes visual features, and has stronger interpretability. Specifically, the framework extracts the source domain, target domain, and metaphor interpretation from metaphorical sentences, separating the elements of the metaphor to deepen the understanding of its themes and intentions. Additionally, the framework introduces image data from the source domain to capture visual similarities and generate visual enhancement prompts specific to the domain. Finally, these prompts are combined with metaphorical interpretation sentences to form the final prompt text. Experimental results demonstrate that this approach effectively captures the essence of metaphor and generates metaphorical images consistent with the textual meaning.

Publisher

Springer Science and Business Media LLC

Link

https://link.springer.com/content/pdf/10.1007/s11063-024-11609-w.pdf

Reference30 articles.

1. Hessel J, Marasović A, Hwang JD, Lee L, Da J, Zellers R, Mankoff R, Choi Y (2023) Do androids laugh at electric sheep? Humor “understanding” benchmarks from the new yorker caption contest

2. Yuri B, Simon D (2020) Sky + fire = sunset. exploring parallels between visually grounded metaphors and image classifiers. In: Beigman KB, Ekaterina S, Patricia L, Smaranda M, Chee W, Anna F, Debanjan G (eds) Proceedings of the second workshop on figurative language processing, pp 126–135, Online. Association for Computational Linguistics

3. Robin R, Andreas B, Dominik L, Patrick E, Björn O (2022) High-resolution image synthesis with latent diffusion models. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp 10684–10695

4. Aditya R, Mikhail P, Gabriel G, Scott G, Chelsea V, Alec R, Mark C, Ilya S (2021) Zero-shot text-to-image generation. In: International conference on machine learning, pp 8821–8831. PMLR

5. Alex N, Prafulla D, Aditya R, Pranav S, Pamela M, Bob M, Ilya S, Mark C (2022) Glide: Towards photorealistic image generation and editing with text-guided diffusion models arxiv:2205.13168v1