Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
110 tokens/sec
GPT-4o
56 tokens/sec
Gemini 2.5 Pro Pro
44 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

On monoaural speech enhancement for automatic recognition of real noisy speech using mixture invariant training (2205.01751v2)

Published 3 May 2022 in cs.SD, cs.CL, and eess.AS

Abstract: In this paper, we explore an improved framework to train a monoaural neural enhancement model for robust speech recognition. The designed training framework extends the existing mixture invariant training criterion to exploit both unpaired clean speech and real noisy data. It is found that the unpaired clean speech is crucial to improve quality of separated speech from real noisy speech. The proposed method also performs remixing of processed and unprocessed signals to alleviate the processing artifacts. Experiments on the single-channel CHiME-3 real test sets show that the proposed method improves significantly in terms of speech recognition performance over the enhancement system trained either on the mismatched simulated data in a supervised fashion or on the matched real data in an unsupervised fashion. Between 16% and 39% relative WER reduction has been achieved by the proposed system compared to the unprocessed signal using end-to-end and hybrid acoustic models without retraining on distorted data.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (4)
  1. Jisi Zhang (9 papers)
  2. Rama Doddipatla (28 papers)
  3. Jon Barker (26 papers)
  4. Catalin Zorila (11 papers)
Citations (4)

Summary

We haven't generated a summary for this paper yet.