Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
110 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

VISinger2+: End-to-End Singing Voice Synthesis Augmented by Self-Supervised Learning Representation (2406.08761v2)

Published 13 Jun 2024 in cs.SD and eess.AS

Abstract: Singing Voice Synthesis (SVS) has witnessed significant advancements with the advent of deep learning techniques. However, a significant challenge in SVS is the scarcity of labeled singing voice data, which limits the effectiveness of supervised learning methods. In response to this challenge, this paper introduces a novel approach to enhance the quality of SVS by leveraging unlabeled data from pre-trained self-supervised learning models. Building upon the existing VISinger2 framework, this study integrates additional spectral feature information into the system to enhance its performance. The integration aims to harness the rich acoustic features from the pre-trained models, thereby enriching the synthesis and yielding a more natural and expressive singing voice. Experimental results in various corpora demonstrate the efficacy of this approach in improving the overall quality of synthesized singing voices in both objective and subjective metrics.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (5)
  1. Yifeng Yu (34 papers)
  2. Jiatong Shi (82 papers)
  3. Yuning Wu (20 papers)
  4. Shinji Watanabe (416 papers)
  5. Yuxun Tang (13 papers)
Citations (2)

Summary

We haven't generated a summary for this paper yet.