Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
5 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Spoken Language Corpora Augmentation with Domain-Specific Voice-Cloned Speech (2406.07090v2)

Published 11 Jun 2024 in eess.AS

Abstract: In this paper we study the impact of augmenting spoken language corpora with domain-specific synthetic samples for the purpose of training a speech recognition system. Using both a conventional neural TTS system and a zero-shot one with voice cloning ability we generate speech corpora that vary in the number of voices. We compare speech recognition models trained with addition of different amounts of synthetic data generated using these two methods with a baseline model trained solely on voice recordings. We show that while the quality of voice-cloned dataset is lower, its increased multivoiceity makes it much more effective than the one with only a few voices synthesized with the use of a conventional neural TTS system. Furthermore, our experiments indicate that using low variability synthetic speech quickly leads to saturation in the quality of the ASR whereas high variability speech provides improvement even when increasing total amount of data used for training by 30%.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (7)
  1. Mateusz Czyżnikiewicz (2 papers)
  2. Łukasz Bondaruk (3 papers)
  3. Jakub Kubiak (3 papers)
  4. Łukasz Degórski (1 paper)
  5. Marek Kubis (8 papers)
  6. Paweł Skórzewski (3 papers)
  7. Adam Wiącek (1 paper)
Citations (1)

Summary

We haven't generated a summary for this paper yet.