Skip to main navigation Skip to search Skip to main content

Active Learning for Text-to-Speech Synthesis with Informative Sample Collection

Research output: Chapter in Book/Report/Conference proceedingConference contribution

Abstract

The construction of high-quality datasets is a cornerstone of modern text-to-speech (TTS) systems. However, the increasing scale of available data poses significant challenges, including storage constraints. To address these issues, we propose a TTS corpus construction method based on active learning. Unlike traditional feed-forward and model-agnostic corpus construction approaches, our method iteratively alternates between data collection and model training, thereby focusing on acquiring data that is more informative for model improvement. This approach enables the construction of a data-efficient corpus. Experimental results demonstrate that the corpus constructed using our method enables higher-quality speech synthesis than corpora of the same size.

Original languageEnglish
Title of host publication2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
PublisherInstitute of Electrical and Electronics Engineers Inc.
Pages903-908
Number of pages6
ISBN (Electronic)9798331572068
DOIs
Publication statusPublished - 2025
Event17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025 - Singapore, Singapore
Duration: 2025 Oct 222025 Oct 24

Publication series

Name2025 Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025

Conference

Conference17th Asia Pacific Signal and Information Processing Association Annual Summit and Conference, APSIPA ASC 2025
Country/TerritorySingapore
CitySingapore
Period25/10/2225/10/24

ASJC Scopus subject areas

  • Artificial Intelligence
  • Computer Science Applications
  • Hardware and Architecture
  • Signal Processing

Fingerprint

Dive into the research topics of 'Active Learning for Text-to-Speech Synthesis with Informative Sample Collection'. Together they form a unique fingerprint.

Cite this