about
Data Collection and Labeling Techniques for Machine Learning (arxiv.org)
2 points by PaulHoule on Aug 8, 2024 | hide | past | pdf | discuss on HN

In plain words: A survey of ways to gather training examples and tag them with correct answers, plus tricks for improving data and models already in hand. It combines machine learning and database perspectives to map what works today and where the open problems lie.

Abstract

Data collection and labeling are critical bottlenecks in the deployment of machine learning applications. With the increasing complexity and diversity of applications, the need for efficient and scalable data collection and labeling techniques has become paramount. This paper provides a review of the state-of-the-art methods in data collection, data labeling, and the improvement of existing data and models. By integrating perspectives from both the machine learning and data management communities, we aim to provide a holistic view of the current landscape and identify future research directions.

Qianyu Huang, Tongfang Zhao
arXiv:2407.12793 · cs.DB, cs.AI, cs.LG · submitted Jun 19, 2024
abstract · pdf · html

add comment on HN