Towards Visually Grounded Sub-Word Speech Unit Discovery (1902.08213v1)

Published 21 Feb 2019 in cs.CL, cs.LG, cs.SD, and eess.AS

Abstract: In this paper, we investigate the manner in which interpretable sub-word speech units emerge within a convolutional neural network model trained to associate raw speech waveforms with semantically related natural image scenes. We show how diphone boundaries can be superficially extracted from the activation patterns of intermediate layers of the model, suggesting that the model may be leveraging these events for the purpose of word recognition. We present a series of experiments investigating the information encoded by these events.

Citations (35)

View on Semantic Scholar

Summary

We haven't generated a summary for this paper yet.

Summarize Now

Towards Visually Grounded Sub-Word Speech Unit Discovery (1902.08213v1)

Summary

Related Papers