A copy of this work was available on the public web and has been preserved in the Wayback Machine. The capture dates from 2020; you can also visit <a rel="external noopener" href="https://arxiv.org/pdf/1611.06986v1.pdf">the original URL</a>. The file type is <code>application/pdf</code>.
Robust end-to-end deep audiovisual speech recognition
[article]
<span title="2016-11-21">2016</span>
<i >
arXiv
</i>
<span class="release-stage" >pre-print</span>
Speech is one of the most effective ways of communication among humans. Even though audio is the most common way of transmitting speech, very important information can be found in other modalities, such as vision. Vision is particularly useful when the acoustic signal is corrupted. Multi-modal speech recognition however has not yet found wide-spread use, mostly because the temporal alignment and fusion of the different information sources is challenging. This paper presents an end-to-end
<span class="external-identifiers">
<a target="_blank" rel="external noopener" href="https://arxiv.org/abs/1611.06986v1">arXiv:1611.06986v1</a>
<a target="_blank" rel="external noopener" href="https://fatcat.wiki/release/yqhvju5jirflbcgirpvaegzpwe">fatcat:yqhvju5jirflbcgirpvaegzpwe</a>
</span>
more »
... sual speech recognizer (AVSR), based on recurrent neural networks (RNN) with a connectionist temporal classification (CTC) loss function. CTC creates sparse "peaky" output activations, and we analyze the differences in the alignments of output targets (phonemes or visemes) between audio-only, video-only, and audio-visual feature representations. We present the first such experiments on the large vocabulary IBM ViaVoice database, which outperform previously published approaches on phone accuracy in clean and noisy conditions.
<a target="_blank" rel="noopener" href="https://web.archive.org/web/20200929160852/https://arxiv.org/pdf/1611.06986v1.pdf" title="fulltext PDF download" data-goatcounter-click="serp-fulltext" data-goatcounter-title="serp-fulltext">
<button class="ui simple right pointing dropdown compact black labeled icon button serp-button">
<i class="icon ia-icon"></i>
Web Archive
[PDF]
<div class="menu fulltext-thumbnail">
<img src="https://blobs.fatcat.wiki/thumbnail/pdf/f7/04/f704e6f70729edeb1ff0087e9a87d1892f9be598.180px.jpg" alt="fulltext thumbnail" loading="lazy">
</div>
</button>
</a>
<a target="_blank" rel="external noopener" href="https://arxiv.org/abs/1611.06986v1" title="arxiv.org access">
<button class="ui compact blue labeled icon button serp-button">
<i class="file alternate outline icon"></i>
arxiv.org
</button>
</a>