Goal-Directed Exploration for Learning Vowels and Syllables: A Computational Model of Speech Acquisition-Reference-Cited by-同舟云学术

Goal-Directed Exploration for Learning Vowels and Syllables: A Computational Model of Speech Acquisition

Published:2021-01-27 Issue:1 Volume:35 Page:53-70
ISSN:0933-1875
Container-title:KI - Künstliche Intelligenz
language:en
Short-container-title:Künstl Intell

Author:

Philippsen Anja^ORCID

Abstract

AbstractInfants learn to speak rapidly during their first years of life, gradually improving from simple vowel-like sounds to larger consonant-vowel complexes. Learning to control their vocal tract in order to produce meaningful speech sounds is a complex process which requires to learn the relationship between motor and sensory processes. In this paper, a computational framework is proposed that models the problem of learning articulatory control for a physiologically plausible 3-D vocal tract model using a developmentally-inspired approach. The system babbles and explores efficiently in a low-dimensional space of goals that are relevant to the learner in its synthetic environment. The learning process is goal-directed and self-organized, and yields an inverse model of the mapping between sensory space and motor commands. This study provides a unified framework that can be used for learning static as well as dynamic motor representations. The successful learning of vowel and syllable sounds as well as the benefit of active and adaptive learning strategies are demonstrated. Categorical perception is found in the acquired models, suggesting that the framework has the potential to replicate phenomena of human speech acquisition.

Funder

Deutsche Forschungsgemeinschaft

Publisher

Springer Science and Business Media LLC

Subject

Artificial Intelligence

Link

http://link.springer.com/content/pdf/10.1007/s13218-021-00704-y.pdf

Reference80 articles.

1. Vouloumanos A, Werker JF (2004) Tuned to the signal: the privileged status of speech for young infants. Dev Sci 7(3):270–276

2. Werker JF, Yeung HH (2005) Infant speech perception bootstraps word learning. Trends Cogn Sci 9(11):519–527

3. Hannun A, Case C, Casper J, Catanzaro B, Diamos G, Elsen E, Prenger R, Satheesh S, Sengupta S, Coates A et al (2014) Deep speech: Scaling up end-to-end speech recognition. arXiv preprint. arXiv:14125567

4. Pratap V, Hannun A, Xu Q, Cai J, Kahn J, Synnaeve G, Liptchinsky V, Collobert R (2019) Wav2letter++: a fast open-source speech recognition system. In: ICASSP 2019–2019 IEEE international conference on acoustics. Speech and signal processing (ICASSP), IEEE, pp 6460–6464

5. Xiong W, Droppo J, Huang X, Seide F, Seltzer M, Stolcke A, Yu D, Zweig G (2016) Achieving human parity in conversational speech recognition. arXiv preprint. arXiv:161005256

Cited by 8 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. Artificial vocal learning guided by speech recognition: What it may tell us about how children learn to speak;Journal of Phonetics;2024-07

2. Simulating vocal learning of spoken language: Beyond imitation;Speech Communication;2023-02

3. Artificial Vocal Learning Guided by Phoneme Recognition and Visual Information;IEEE/ACM Transactions on Audio, Speech, and Language Processing;2023

4. Exploration strategies for articulatory synthesis of complex syllable onsets;Interspeech 2022;2022-09-18

5. Application of Big Data Mining Technology in the Digital Construction of Vocal Music Teaching Resource Library;Wireless Communications and Mobile Computing;2022-07-29