メインナビゲーションにスキップ 検索にスキップ メインコンテンツにスキップ

Improving Speech Prosody of Audiobook Text-To-Speech Synthesis with Acoustic and Textual Contexts

  • Detai Xin
  • , Sharath Adavanne
  • , Federico Ang
  • , Ashish Kulkarni
  • , Shinnosuke Takamichi
  • , Hiroshi Saruwatari

研究成果: Conference contribution

抄録

We present a multi-speaker Japanese audiobook text-to-speech (TTS) system that leverages multimodal context information of preceding acoustic context and bilateral textual context to improve the prosody of synthetic speech. Previous work either uses unilateral or single-modality context, which does not fully represent the context information. The proposed method uses an acoustic context encoder and a textual context encoder to aggregate context information and feeds it to the TTS model, which enables the model to predict context-dependent prosody. We conducted comprehensive objective and subjective evaluations on a multi-speaker Japanese audiobook dataset. Experimental results demonstrate that the proposed method significantly outperforms two previous works. Additionally, we present insights about the different choices of context - modalities, lateral information and length - for audiobook TTS that have never been discussed in the literature before.

本文言語English
ホスト出版物のタイトルICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing, Proceedings
出版社Institute of Electrical and Electronics Engineers Inc.
ISBN(電子版)9781728163277
DOI
出版ステータスPublished - 2023
外部発表はい
イベント48th IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2023 - Rhodes Island, Greece
継続期間: 2023 6月 42023 6月 10

出版物シリーズ

名前ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
2023-June
ISSN(印刷版)1520-6149

Conference

Conference48th IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2023
国/地域Greece
CityRhodes Island
Period23/6/423/6/10

ASJC Scopus subject areas

  • ソフトウェア
  • 信号処理
  • 電子工学および電気工学

フィンガープリント

「Improving Speech Prosody of Audiobook Text-To-Speech Synthesis with Acoustic and Textual Contexts」の研究トピックを掘り下げます。これらがまとまってユニークなフィンガープリントを構成します。

引用スタイル