Parents and Children: Distinguishing Multimodal DeepFakes from Natural Images-Reference-Cited by-同舟云学术

Parents and Children: Distinguishing Multimodal DeepFakes from Natural Images

Published:2024-05-21 Issue: Volume: Page:
ISSN:1551-6857
Container-title:ACM Transactions on Multimedia Computing, Communications, and Applications
language:en
Short-container-title:ACM Trans. Multimedia Comput. Commun. Appl.

Author:

Amoroso Roberto¹^ORCID,Morelli Davide²^ORCID,Cornia Marcella¹^ORCID,Baraldi Lorenzo¹^ORCID,Del Bimbo Alberto³^ORCID,Cucchiara Rita⁴^ORCID

Affiliation:

1. University of Modena and Reggio Emilia, Italy

2. University of Modena and Reggio Emilia, Italy and University of Pisa, Italy

3. University of Florence, Italy

4. University of Modena and Reggio Emilia, Italy and IIT-CNR, Italy

Abstract

Recent advancements in diffusion models have enabled the generation of realistic deepfakes from textual prompts in natural language. While these models have numerous benefits across various sectors, they have also raised concerns about the potential misuse of fake images and cast new pressures on fake image detection. In this work, we pioneer a systematic study on deepfake detection generated by state-of-the-art diffusion models. Firstly, we conduct a comprehensive analysis of the performance of contrastive and classification-based visual features, respectively extracted from CLIP-based models and ResNet or ViT-based architectures trained on image classification datasets. Our results demonstrate that fake images share common low-level cues, which render them easily recognizable. Further, we devise a multimodal setting wherein fake images are synthesized by different textual captions, which are used as seeds for a generator. Under this setting, we quantify the performance of fake detection strategies and introduce a contrastive-based disentangling method that lets us analyze the role of the semantics of textual descriptions and low-level perceptual cues. Finally, we release a new dataset, called COCOFake, containing about 1.2M images generated from the original COCO image-caption pairs using two recent text-to-image diffusion models, namely Stable Diffusion v1.4 and v2.0.

Publisher

Association for Computing Machinery (ACM)

Link

https://dl.acm.org/doi/pdf/10.1145/3665497

Reference75 articles.

1. Manuele Barraco, Sara Sarto, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2023. With a Little Help from your own Past: Prototypical Memory Networks for Image Captioning. In Proceedings of the IEEE/CVF International Conference on Computer Vision.

2. Miles Brundage Shahar Avin Jack Clark Helen Toner Peter Eckersley Ben Garfinkel Allan Dafoe Paul Scharre Thomas Zeitzoff Bobby Filar et al. 2018. The Malicious Use of Artificial Intelligence: Forecasting Prevention and Mitigation. arXiv preprint arXiv:1802.07228 (2018).

3. A limited memory algorithm for bound constrained optimization;Byrd Richard H;SIAM Journal on Scientific Computing,1995

4. Davide Caffagni, Manuele Barraco, Marcella Cornia, Lorenzo Baraldi, and Rita Cucchiara. 2023. SynthCap: Augmenting Transformers with Synthetic Data for Image Captioning. In Proceedings of the International Conference on Image Analysis and Processing.

5. Lucy Chai, David Bau, Ser-Nam Lim, and Phillip Isola. 2020. What Makes Fake Images Detectable? Understanding Properties that Generalize. In Proceedings of the European Conference on Computer Vision.

Cited by 1 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Detecting images generated by diffusers;PeerJ Computer Science;2024-07-10