about
All the World's a (Hyper)Graph: A Data Drama (arxiv.org)
3 points by Topolomancer on Jun 22, 2022 | hide | past | pdf | discuss on HN

In plain words: Shakespeare's plays are turned into many versions of relationship data, from simple character-meeting graphs to richer group structures, so anyone can test how much the setup changes results. Many graph mining answers shift with the version chosen, showing today's data-curation habits are shaky.

Abstract

We introduce Hyperbard, a dataset of diverse relational data representations derived from Shakespeare's plays. Our representations range from simple graphs capturing character co-occurrence in single scenes to hypergraphs encoding complex communication settings and character contributions as hyperedges with edge-specific node weights. By making multiple intuitive representations readily available for experimentation, we facilitate rigorous representation robustness checks in graph learning, graph mining, and network analysis, highlighting the advantages and drawbacks of specific representations. Leveraging the data released in Hyperbard, we demonstrate that many solutions to popular graph mining problems are highly dependent on the representation choice, thus calling current graph curation practices into question. As an homage to our data source, and asserting that science can also be art, we present all our points in the form of a play.

Corinna Coupette, Jilles Vreeken, Bastian Rieck
arXiv:2206.08225 · cs.LG, cs.CL, cs.CY, cs.SI · submitted Jun 16, 2022 · updated Dec 6, 2023
abstract · pdf · html · This is the full version of our paper; an abridged version appears in Digital Scholarship in the Humanities. Landing page for code and data: https://hyperbard.net/

add comment on HN
Also discussed: Jun 2022 (1 point, 0 comments)