Skip to main navigation Skip to search Skip to main content

Improving Speech Prosody of Audiobook Text-To-Speech Synthesis with Acoustic and Textual Contexts

  • Detai Xin
  • , Sharath Adavanne
  • , Federico Ang
  • , Ashish Kulkarni
  • , Shinnosuke Takamichi
  • , Hiroshi Saruwatari

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

We present a multi-speaker Japanese audiobook text-to-speech (TTS) system that leverages multimodal context information of preceding acoustic context and bilateral textual context to improve the prosody of synthetic speech. Previous work either uses unilateral or single-modality context, which does not fully represent the context information. The proposed method uses an acoustic context encoder and a textual context encoder to aggregate context information and feeds it to the TTS model, which enables the model to predict context-dependent prosody. We conducted comprehensive objective and subjective evaluations on a multi-speaker Japanese audiobook dataset. Experimental results demonstrate that the proposed method significantly outperforms two previous works. Additionally, we present insights about the different choices of context - modalities, lateral information and length - for audiobook TTS that have never been discussed in the literature before.

Original languageEnglish
Title of host publicationICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing, Proceedings
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9781728163277
DOIs
Publication statusPublished - 2023
Externally publishedYes
Event48th IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2023 - Rhodes Island, Greece
Duration: 2023 Jun 42023 Jun 10

Publication series

NameICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
Volume2023-June
ISSN (Print)1520-6149

Conference

Conference48th IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2023
Country/TerritoryGreece
CityRhodes Island
Period23/6/423/6/10

Keywords

  • TTS
  • audiobook
  • context modeling
  • speech prosody
  • text-to-speech synthesis

ASJC Scopus subject areas

  • Software
  • Signal Processing
  • Electrical and Electronic Engineering

Fingerprint

Dive into the research topics of 'Improving Speech Prosody of Audiobook Text-To-Speech Synthesis with Acoustic and Textual Contexts'. Together they form a unique fingerprint.

Cite this