about
Auto-Labeling Data for Object Detection (arxiv.org)
6 points by nwlotz on Jun 4, 2025 | hide | past | pdf | 1 comment on HN

In plain words: A system uses models that already link images and text to write object labels for training pictures, so no people have to label them. The small detectors trained on these labels kept accuracy close to human-labeled ones on several datasets while cutting labeling time and cost.

Abstract

Great labels make great models. However, traditional labeling approaches for tasks like object detection have substantial costs at scale. Furthermore, alternatives to fully-supervised object detection either lose functionality or require larger models with prohibitive computational costs for inference at scale. To that end, this paper addresses the problem of training standard object detection models without any ground truth labels. Instead, we configure previously-trained vision-language foundation models to generate application-specific pseudo "ground truth" labels. These auto-generated labels directly integrate with existing model training frameworks, and we subsequently train lightweight detection models that are computationally efficient. In this way, we avoid the costs of traditional labeling, leverage the knowledge of vision-language models, and keep the efficiency of lightweight models for practical application. We perform exhaustive experiments across multiple labeling configurations, downstream inference models, and datasets to establish best practices and set an extensive auto-labeling benchmark. From our results, we find that our approach is a viable alternative to standard labeling in that it maintains competitive performance on multiple datasets and substantially reduces labeling time and costs.

Brent A. Griffin, Manushree Gangwar, Jacob Sela, Jason J. Corso
arXiv:2506.02359 · cs.CV · submitted Jun 3, 2025
abstract · pdf · html

add comment on HN

The ML research team at Voxel51 just released a paper showing that foundation models rival the accuracy of human annotators in labeling large visual datasets, at several orders of magnitude less time and cost.

We also found that models trained from these labels perform about as well as those trained from human labels when tested against public validation sets. Interestingly, setting a relatively low confidence threshold (0.2 - 0.5) for the auto-generated labels maximized downstream model performance. Very high confidence thresholds often produced worse results due to reduced recall.

The upshot is that zero-shot labeling can replace human annotation in many datasets. The massive cost savings can then be redirected toward training higher-parameter models.

Happy to answer any questions about the research. You can also read this blog we wrote that goes more in depth into the methods and tools we used. https://link.voxel51.com/HN-VAL-blog/