Artfusion: A Diffusion Model-Based Style Synthesis Framework for Portraits-Reference-Cited by-同舟云学术

Artfusion: A Diffusion Model-Based Style Synthesis Framework for Portraits

Published:2024-01-25 Issue:3 Volume:13 Page:509
ISSN:2079-9292
Container-title:Electronics
language:en
Short-container-title:Electronics

Author:

Yang Hyemin¹,Yang Heekyung²^ORCID,Min Kyungha¹

Affiliation:

1. Department of Computer Science, Sangmyung University, Seoul 03016, Republic of Korea

2. Department of Software, Sangmyung University, Cheonan 31066, Republic of Korea

Abstract

We present a diffusion model-based approach that applies the artistic style of an artist or an art movement to a portrait photograph. Learning the style from the artworks of an artist or an art movement requires a training dataset composed of a lot of samples. We resolve this limitation by combining Contrastive Language Image Pretraining (CLIP) encoder and diffusion model, since the CLIP encoder extracts the features from an input portrait in a very effective way. Our framework includes three independent CLIP encoders that extract the text features, color features and Canny edge features from an input portrait, respectively. These features are incorporated to the style information extracted through a diffusion model to complete the stylization on an input portrait. The diffusion model extracts the style information from the sample images in the training dataset using an image encoder. The denoising steps in the diffusion model applies the style information from the training dataset to the CLIP-based features from an input portrait. Finally, our framework produces an artistic portrait that presents both the identity of the input portrait and the artistic style from the training dataset. The most important contribution of our framework is that our framework requires less than a hundred sample images for an artistic style. Therefore, our framework can successfully extract styles from an artist who has drawn less than a hundred artworks. We sample three artists and three art movements and apply these styles to the portraits of various identities and produce visually pleasing results. We evaluate our results using various metrics, including Frechet Inception Distance (FID), ArtFID and Language-Image Quality Evaluator (LIQE) to prove the excellence of our results.

Funder

Ministry of Education

Publisher

MDPI AG

Link

https://www.mdpi.com/2079-9292/13/3/509/pdf

Reference37 articles.

1. Gatys, L.A., Ecker, A.S., and Bethge, M. (July, January 26). Image style transfer using convolutional neural networks. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Las Vegas, NV, USA.

2. Huang, X., and Benlongie, S. (2017, January 21–26). Arbitrary style transfer in real-time with adaptive instance Normalization. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.

3. An, J., Huang, S., Song, Y., Dou, D., Liu, W., and Luo, J. (2017, January 21–26). ArtFlow: Unbiased image style transfer via reversible neural flows. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Honolulu, HI, USA.

4. Deng, Y., Tang, F., Dong, W., Sun, W., Huang, F., and Luo, J. (2020, January 12–16). Arbitrary style transfer via multi-adaptation network. Proceedings of the 28th ACM International Conference on Multimedia, Seattle, WA, USA.

5. Chen, H., Wang, Z., Zhang, H., Zuo, Z., Li, A., Xing, W., and Lu, D. (2021, January 6–14). Artistic style transfer with internal-externel learning and contrastive learning. Proceedings of the Conference on Neural Information Processing Systems, Online.