Procedurally generated AI compound media for expanding audial creations, broadening immersion and perception experience-Reference-Cited by-同舟云学术

Procedurally generated AI compound media for expanding audial creations, broadening immersion and perception experience

Published:2024-06-25 Issue: Volume: Page:341-348
ISSN:2300-1933
Container-title:International Journal of Electronics and Telecommunications
language:pl
Short-container-title:

Author:

Samson Grzegorz¹

Affiliation:

1. Feliks Nowowiejski Academy of Music in Bydgoszcz, Poland

Abstract

Recently, the world has been gaining vastly increasing access to more and more advanced artificial intelligence tools. This phenomenon does not bypass the world of sound and visual art, and both of these worlds can benefit in ways yet unexplored, drawing them closer to one another. Recent breakthroughs open possibilities to utilize AI driven tools for creating generative art and using it as a compound of other multimedia. The aim of this paper is to present an original concept of using AI to create a visual compound material to existing audio source. This is a way of broadening accessibility thus appealing to different human senses using source media, expanding its initial form. This research utilizes a novel method of enhancing fundamental material consisting of text audio or text source (script) and sound layer (audio play) by adding an extra layer of multimedia experience – a visual one, generated procedurally. A set of images generated by AI tools, creating a story-telling animation as a new way to immerse into the experience of sound perception and focus on the initial audial material. The main idea of the paper consists of creating a pipeline, form of a blueprint for the process of procedural image generation based on the source context (audial or textual) transformed into text prompts and providing tools to automate it by programming a set of code instructions. This process allows creation of coherent and cohesive (to a certain extent) visual cues accompanying audial experience levering it to multimodal piece of art. Using nowadays technologies, creators can enhance audial forms procedurally, providing them with visual context. The paper refers to current possibilities, use cases, limitations and biases giving presented tools and solutions.

Publisher

Polish Academy of Sciences Chancellery

Link

https://journals.pan.pl/Content/131791/10_4596_Samson_sk.pdf