Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
5 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

DCCRGAN: Deep Complex Convolution Recurrent Generator Adversarial Network for Speech Enhancement (2012.10732v2)

Published 19 Dec 2020 in eess.AS and cs.SD

Abstract: Generative adversarial network (GAN) still exists some problems in dealing with speech enhancement (SE) task. Some GAN-based systems adopt the same structure from Pixel-to-Pixel directly without special optimization. The importance of the generator network has not been fully explored. Other related researches change the generator network but operate in the time-frequency domain, which ignores the phase mismatch problem. In order to solve these problems, a deep complex convolution recurrent GAN (DCCRGAN) structure is proposed in this paper. The complex module builds the correlation between magnitude and phase of the waveform and has been proved to be effective. The proposed structure is trained in an end-to-end way. Different LSTM layers are used in the generator network to sufficiently explore the speech enhancement performance of DCCRGAN. The experimental results confirm that the proposed DCCRGAN outperforms the state-of-the-art GAN-based SE systems.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (5)
  1. Huixiang Huang (1 paper)
  2. Renjie Wu (8 papers)
  3. Jingbiao Huang (1 paper)
  4. Jucai Lin (2 papers)
  5. Jun Yin (108 papers)
Citations (6)

Summary

We haven't generated a summary for this paper yet.