about
End-To-end Audiovisual Speech Recognition (arxiv.org)
54 points by ghosthamlet on Mar 8, 2018 | hide | past | pdf | 12 comments on HN

In plain words: A system reads mouth movements from video frames and raw sound waves together, then combines both to recognize whole spoken words. It slightly beat audio-only systems in quiet conditions but far outperformed them when the audio was noisy.

Abstract · End-to-end Audiovisual Speech Recognition

Several end-to-end deep learning approaches have been recently presented which extract either audio or visual features from the input images or audio signals and perform speech recognition. However, research on end-to-end audiovisual models is very limited. In this work, we present an end-to-end audiovisual model based on residual networks and Bidirectional Gated Recurrent Units (BGRUs). To the best of our knowledge, this is the first audiovisual fusion model which simultaneously learns to extract features directly from the image pixels and audio waveforms and performs within-context word recognition on a large publicly available dataset (LRW). The model consists of two streams, one for each modality, which extract features directly from mouth regions and raw waveforms. The temporal dynamics in each stream/modality are modeled by a 2-layer BGRU and the fusion of multiple streams/modalities takes place via another 2-layer BGRU. A slight improvement in the classification rate over an end-to-end audio-only and MFCC-based model is reported in clean audio conditions and low levels of noise. In presence of high levels of noise, the end-to-end audiovisual model significantly outperforms both audio-only models.

Stavros Petridis, Themos Stafylakis, Pingchuan Ma, Feipeng Cai, Georgios Tzimiropoulos, Maja Pantic
arXiv:1802.06424 · cs.CV · submitted Feb 18, 2018 · updated Feb 22, 2018
abstract · pdf · html · Accepted to ICASSP 2018

add comment on HN

Y'all didn't cite some extremely relevant work: https://arxiv.org/abs/1709.00572 (disclaimer - they're from my lab)
What is HNs impression of the etiquette of "demanding" citations? I agree with nmca that it's the author's onus to research prior work as widely as possible, but if something is missed, presumably both by the authors and the reviewers, and therefore clearly was not actually a reference for the work in question, is the author really at fault? Is it "ok" to demand a citation like this?
If a primary source shows up on HN with a link to relevant research from their lab, that's great! That adds wonderful value to HN. I don't think we need to read nmca's comment as a demand.
Closed-form peer review is basically a (very tight) spam filter in ML at the moment. Everyone wants the papers to appear quickly, so reviewers can't meaningfully require changes --- papers are in or out. Overall this is net better. Requiring changes leads to bike-shedding, and really long publication processes.

Most review is community review, of exactly this sort of form. So we really don't want to add a lot of politeness constraints around requesting citations.

The literature is moving so fast that authors can get away with assuming readers won't know about relevant work. If we let that thrive, we'll reward bad actors who "forget" prior work and claim novelty.

>> Is it "ok" to demand a citation like this?

Sure, why not

(especially with disclaimer given)

There is nothing wrong with asking for a citation to a relevant work, especially if you explain why the work is relevant. The authors still have no obligation to cite you.

But demanding a citation by posting here, on a link aggregation site, seems more like a self advertisement, and doesn't seem appropriate to me. Who knows if the authors will even see this? Send them an email.

Just for the record I think all of this is very reasonable, with a couple of minor caveats. It's not a self advertisement, just an ad for some colleagues. Also I think the discussion above is interesting, but operates on the presumption that I demanded a citation. I didn't, I just said they didn't cite something that is (implicitly in my opinion) very relevant. Indeed, it is not always possible or appropriate to cite everything that's very relevant, and their (almost certainly more valid) opinion of what's relevant might vary.

Also this seemed like a fun and bizarrely civil discussion for the internet. Well done all.

> Also this seemed like a fun and bizarrely civil discussion for the internet. Well done all.

:) Just for the record my question was posed as it was because it was in earnest. I actually don't know what's correct. I do agree that the authors should cite as much as possible, yet I react in a weird way when I see people say, why didn't you cite X, Y, Z.. because in one sense citations are necessary for credit attribution, but on the other hand if you truly did something in parallel, without knowledge of another source, does it make sense to cite it, since it's clearly not necessary to explain your work? I suppose it does, since the civility of giving credit where due is part of science, but it also feels a little unfair to behold authors to include similar works that they don't have time to re-implement and compare against for example, and cause them to mention with excuses like, "I just found out about this one 5 minutes ago", or even cause them to delay publication in order to include a brand-new benchmark that was just published a few days before their deadline.

Anyways, I think I only asked the question in response to your statement because of the way you posed it, "Y'all didn't cite..", rather than, "You might want to check out for future work.." I do apologize if I interpreted that differently than you intended, but I do think that "didn't cite" does come off more as an accusation than, "you might not be aware of".

Re-reading my first paragraph, actually I think it's not so out of the norm to say, "we didn't compare with X,Y,Z because we were not aware of those works sufficiently before the deadline, having been very recently published." It's a different story if the relevant work is like a year or more old I guess.

Wow: "In presence of high levels of noise, the end-to-end audiovisual model significantly outperforms both audio-only models."
I am curious what the success rate would be without any audio - by just reading lips and face.
I'm going to guess "very bad" given that even experts can't do this very well. I don't think there is enough information.

Edit: It does say in the paper actually. The "classification rate" (I assume % of words correctly identified - they only recognise single works from a dictionary of 500) is 82 for visual only and 98 for audio, and audio-visual (without noise).

Without noise the audio is good enough to be near perfect (I assume 100 is perfect) so the video doesn't really help. It helps when there is noise (which matches real life experience - people lip read in noisy bars).

> people lip read in noisy bars

Maybe that explains why I, as a visually impaired person who can't lip read, have trouble carrying on conversations in noisy bars. Good to know!