In plain words: A system watches the game footage and reads the live Twitch chat, slang and all in English or Chinese, to find the exciting moments worth showing. Tested on League of Legends championship streams, it picked highlights well using both signals together.
Abstract
Sports channel video portals offer an exciting domain for research on multimodal, multilingual analysis. We present methods addressing the problem of automatic video highlight prediction based on joint visual features and textual analysis of the real-world audience discourse with complex slang, in both English and traditional Chinese. We present a novel dataset based on League of Legends championships recorded from North American and Taiwanese Twitch.tv channels (will be released for further research), and demonstrate strong results on these using multimodal, character-level CNN-RNN model architectures.
Cheng-Yang Fu, Joon Lee, Mohit Bansal, Alexander C. Berg
arXiv:1707.08559 · cs.CL, cs.AI, cs.CV, cs.LG, cs.MM · submitted Jul 26, 2017
abstract · pdf · html · EMNLP 2017