about
Mapping the Podcast Ecosystem with the Structured Podcast Research Corpus (arxiv.org)
2 points by avyfain on Dec 6, 2024 | hide | past | pdf | discuss on HN

In plain words: Built a free collection of transcripts from over 1.1 million English podcast episodes pulled from public feeds in May–June 2020, with audio and speaker details for 370,000 of them. It covers nearly every English podcast reachable this way, far beyond the small samples available before.

Abstract

Podcasts provide highly diverse content to a massive listener base through a unique on-demand modality. However, limited data has prevented large-scale computational analysis of the podcast ecosystem. To fill this gap, we introduce a massive dataset of over 1.1M podcast transcripts that is largely comprehensive of all English language podcasts available through public RSS feeds from May and June of 2020. This data is not limited to text, but rather includes audio features and speaker turns for a subset of 370K episodes, and speaker role inferences and other metadata for all 1.1M episodes. Using this data, we also conduct a foundational investigation into the content, structure, and responsiveness of this ecosystem. Together, our data and analyses open the door to continued computational research of this popular and impactful medium.

Benjamin Litterer, David Jurgens, Dallas Card
arXiv:2411.07892 · cs.CL, cs.CY · submitted Nov 12, 2024 · updated Dec 19, 2025
abstract · pdf · html · 9 pages, 3 figures

add comment on HN