Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
97 tokens/sec
GPT-4o
53 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
4 tokens/sec
GPT-4.1 Pro
47 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

Improving Perceptual Quality, Intelligibility, and Acoustics on VoIP Platforms (2303.09048v1)

Published 16 Mar 2023 in cs.SD, cs.AI, cs.LG, cs.MM, and eess.AS

Abstract: In this paper, we present a method for fine-tuning models trained on the Deep Noise Suppression (DNS) 2020 Challenge to improve their performance on Voice over Internet Protocol (VoIP) applications. Our approach involves adapting the DNS 2020 models to the specific acoustic characteristics of VoIP communications, which includes distortion and artifacts caused by compression, transmission, and platform-specific processing. To this end, we propose a multi-task learning framework for VoIP-DNS that jointly optimizes noise suppression and VoIP-specific acoustics for speech enhancement. We evaluate our approach on a diverse VoIP scenarios and show that it outperforms both industry performance and state-of-the-art methods for speech enhancement on VoIP applications. Our results demonstrate the potential of models trained on DNS-2020 to be improved and tailored to different VoIP platforms using VoIP-DNS, whose findings have important applications in areas such as speech recognition, voice assistants, and telecommunication.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (15)
  1. Joseph Konan (10 papers)
  2. Ojas Bhargave (4 papers)
  3. Shikhar Agnihotri (4 papers)
  4. Hojeong Lee (3 papers)
  5. Ankit Shah (47 papers)
  6. Yunyang Zeng (4 papers)
  7. Amanda Shu (2 papers)
  8. Haohui Liu (2 papers)
  9. Xuankai Chang (61 papers)
  10. Hamza Khalid (3 papers)
  11. Minseon Gwak (3 papers)
  12. Kawon Lee (4 papers)
  13. Minjeong Kim (26 papers)
  14. Bhiksha Raj (180 papers)
  15. Shuo Han (74 papers)
Citations (2)

Summary

We haven't generated a summary for this paper yet.