Four-Features Evaluation of Text to Speech Systems for Three Social Robots-Reference-Cited by-同舟云学术

Four-Features Evaluation of Text to Speech Systems for Three Social Robots

Published:2020-02-05 Issue:2 Volume:9 Page:267
ISSN:2079-9292
Container-title:Electronics
language:en
Short-container-title:Electronics

Author:

Alonso Martin Fernando^ORCID,Malfaz María^ORCID,Castro-González Álvaro^ORCID,Castillo José Carlos^ORCID,Salichs Miguel Ángel^ORCID

Abstract

The success of social robotics is directly linked to their ability of interacting with people. Humans possess verbal and non-verbal communication skills, and, therefore, both are essential for social robots to get a natural human–robot interaction. This work focuses on the first of them since the majority of social robots implement an interaction system endowed with verbal capacities. In order to do this implementation, we must equip social robots with an artificial voice system. In robotics, a Text to Speech (TTS) system is the most common speech synthesizer technique. The performance of a speech synthesizer is mainly evaluated by its similarity to the human voice in relation to its intelligibility and expressiveness. In this paper, we present a comparative study of eight off-the-shelf TTS systems used in social robots. In order to carry out the study, 125 participants evaluated the performance of the following TTS systems: Google, Microsoft, Ivona, Loquendo, Espeak, Pico, AT&T, and Nuance. The evaluation was performed after observing videos where a social robot communicates verbally using one TTS system. The participants completed a questionnaire to rate each TTS system in relation to four features: intelligibility, expressiveness, artificiality, and suitability. In this study, four research questions were posed to determine whether it is possible to present a ranking of TTS systems in relation to each evaluated feature, or, on the contrary, there are no significant differences between them. Our study shows that participants found differences between the TTS systems evaluated in terms of intelligibility, expressiveness, and artificiality. The experiments also indicated that there was a relationship between the physical appearance of the robots (embodiment) and the suitability of TTS systems.

Publisher

MDPI AG

Subject

Electrical and Electronic Engineering,Computer Networks and Communications,Hardware and Architecture,Signal Processing,Control and Systems Engineering

Link

https://www.mdpi.com/2079-9292/9/2/267/pdf

Reference40 articles.

1. Evaluating text-to-speech systems: Some methodological aspects

2. Is text-to-speech synthesis ready for use in computer-assisted language learning?

3. Text-to-speech conversion technology

4. Review of text‐to‐speech conversion for English

5. Top 10 Text to Speech (TTS) Software for eLearning https://elearningindustry.com/top-10-text-to-speech-tts-software-elearning

Cited by 11 articles. 订阅此论文施引文献订阅此论文施引文献，注册后可以免费订阅5篇论文的施引文献，订阅后可以查看论文全部施引文献

1. The symmetric technique of formant transition generation for use in speech synthesis in Arabic;International Journal of Information Technology;2024-07-24

2. Harnessing AI and NLP Tools for Innovating Brand Name Generation and Evaluation: A Comprehensive Review;Multimodal Technologies and Interaction;2024-07-01

3. Giving Robots a Voice: Human-in-the-Loop Voice Creation and open-ended Labeling;Proceedings of the CHI Conference on Human Factors in Computing Systems;2024-05-11

4. Interactive Development of Medical Robot Using Voice Chatbot Based on Deep Learning;2023 29th International Conference on Telecommunications (ICT);2023-11-08

5. Interactive Multimodal Learning: Towards Using Pedagogical Agents for Inclusive Education;2023 IEEE International Humanitarian Technology Conference (IHTC);2023-11-01