about
Drawing Inferences Between Panels in Comic Book Narratives [pdf] (arxiv.org)
1 point by maheshkkumar on Nov 28, 2016 | hide | past | pdf | discuss on HN

In plain words: They collected 1.2 million comic panels with dialogue and made tests where a computer must fill in what happens in a missing panel from the ones before it. Pictures or words alone couldn't follow the plot, and every model tested fell short of humans.

Abstract · The Amazing Mysteries of the Gutter: Drawing Inferences Between Panels in Comic Book Narratives

Visual narrative is often a combination of explicit information and judicious omissions, relying on the viewer to supply missing details. In comics, most movements in time and space are hidden in the "gutters" between panels. To follow the story, readers logically connect panels together by inferring unseen actions through a process called "closure". While computers can now describe what is explicitly depicted in natural images, in this paper we examine whether they can understand the closure-driven narratives conveyed by stylized artwork and dialogue in comic book panels. We construct a dataset, COMICS, that consists of over 1.2 million panels (120 GB) paired with automatic textbox transcriptions. An in-depth analysis of COMICS demonstrates that neither text nor image alone can tell a comic book story, so a computer must understand both modalities to keep up with the plot. We introduce three cloze-style tasks that ask models to predict narrative and character-centric aspects of a panel given n preceding panels as context. Various deep neural architectures underperform human baselines on these tasks, suggesting that COMICS contains fundamental challenges for both vision and language.

Mohit Iyyer, Varun Manjunatha, Anupam Guha, Yogarshi Vyas, Jordan Boyd-Graber, Hal Daumé, Larry Davis
arXiv:1611.05118 · cs.CV, cs.CL · submitted Nov 16, 2016 · updated May 7, 2017
abstract · pdf · html

add comment on HN
Also discussed: Nov 2016 (2 points, 0 comments)