about
FSD50K: An Open Dataset of Human-Labeled Sound Events (arxiv.org)
2 points by sebg on Nov 19, 2020 | hide | past | pdf | discuss on HN

In plain words: A free collection of 51,000 human-labeled sound clips in 200 classes whose audio can be downloaded and shared, unlike AudioSet, which shares only ready-made features and relies on vanishing YouTube videos. Baseline tests showed that how clips are split into training and test sets strongly affects results.

Abstract

Most existing datasets for sound event recognition (SER) are relatively small and/or domain-specific, with the exception of AudioSet, based on over 2M tracks from YouTube videos and encompassing over 500 sound classes. However, AudioSet is not an open dataset as its official release consists of pre-computed audio features. Downloading the original audio tracks can be problematic due to YouTube videos gradually disappearing and usage rights issues. To provide an alternative benchmark dataset and thus foster SER research, we introduce FSD50K, an open dataset containing over 51k audio clips totalling over 100h of audio manually labeled using 200 classes drawn from the AudioSet Ontology. The audio clips are licensed under Creative Commons licenses, making the dataset freely distributable (including waveforms). We provide a detailed description of the FSD50K creation process, tailored to the particularities of Freesound data, including challenges encountered and solutions adopted. We include a comprehensive dataset characterization along with discussion of limitations and key factors to allow its audio-informed usage. Finally, we conduct sound event classification experiments to provide baseline systems as well as insight on the main factors to consider when splitting Freesound audio data for SER. Our goal is to develop a dataset to be widely adopted by the community as a new open benchmark for SER research.

Eduardo Fonseca, Xavier Favory, Jordi Pons, Frederic Font, Xavier Serra
arXiv:2010.00475 · cs.SD, cs.LG, eess.AS, stat.ML · submitted Oct 1, 2020 · updated Apr 23, 2022
abstract · pdf · html · Accepted version in TASLP. Main updates include: estimation of the amount of label noise in FSD50K, SNR comparison between FSD50K and AudioSet, improved description of evaluation metrics including equations, clarification of experimental methodology and some results, some content moved to Appendix for readability. https://ieeexplore.ieee.org/document/9645159

add comment on HN