about
Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models (arxiv.org)
55 points by teleforce 102 days ago | hide | past | pdf | 9 comments on HN

In plain words: A new family of real-time vision models detects objects in one pass, skipping the usual cleanup step that deletes overlapping boxes and trimming the bulky output head. It beats earlier real-time detectors on the accuracy-versus-speed tradeoff, scoring 40.9–57.5 while running in 1.7–11.8 ms.

Abstract

Real-time vision demands models that are accurate, efficient, and simple to deploy across diverse hardware. The YOLO family has become widely deployed for this reason, yet most YOLO detectors still rely on non-maximum suppression at inference, carry heavy detection heads due to Distribution Focal Loss, require long training schedules, and can leave the smallest objects without positive label assignments. We present Ultralytics YOLO26, a unified real-time vision model family that addresses these limitations through coordinated architecture and training advances. YOLO26 uses a dual-head design for native NMS-free end-to-end inference and removes DFL entirely, yielding a lighter head with unconstrained regression range. Its training pipeline combines MuSGD, a hybrid Muon-SGD optimizer adapted from large language model training; Progressive Loss, which shifts supervision toward the inference-time head; and STAL, a label assignment strategy that guarantees positive coverage for small objects. Beyond detection, YOLO26 introduces task-specific head and loss designs for instance segmentation, pose estimation, and oriented detection, producing consistent gains across tasks and scales. The family spans five scales (n/s/m/l/x) and supports detection, instance segmentation, pose estimation, classification, and oriented detection in a single pipeline, with an open-vocabulary extension, YOLOE-26, for text-, visual-, and prompt-free inference. Across all scales, YOLO26 achieves 40.9-57.5 mAP on COCO at 1.7-11.8 ms T4 TensorRT latency, advancing the accuracy-latency Pareto front over prior real-time detectors, while YOLOE-26x reaches 40.6 AP on LVIS minival under text prompting. Code and models are available at https://github.com/ultralytics/ultralytics.

Glenn Jocher, Jing Qiu, Mengyu Liu, Shuai Lyu, Fatih Cagatay Akyon, Muhammet Esat Kalfaoglu
arXiv:2606.03748 · cs.CV, cs.AI · submitted Jun 2, 2026
abstract · pdf · html · 31 pages, 8 figures

add comment on HN

Suspect choice for the paper to only include a single DETR from 2022 in the headline pareto chart and claim to have "the strongest AP–latency trade-off"... Clearly the authors were aware of models that exceed theirs given they even mentioned some of them in the introduction.

> In parallel, DETR [5] cast detection as end-to-end set prediction, and its real-time descendants (RT-DETR [98], D-FINE [55], DEIM [21], RF-DETR [62]) have narrowed the accuracy gap with CNN based detectors on standard benchmarks.

When did YOLO versions jump from 13 to 26? What's next, YOLO XP?
They have switched to release year, not the next available yolo number thing they have been doing.
We already had YOLO9000.
I wait for YOLO vista
My grandma still uses YOLO ME
Lots of attention regarding YOLO today, not sure why. RF-DETR is still SOTA.
Related: An Introduction to YOLO26

77 points, 24 comments

https://news.ycombinator.com/item?id=48639165

it's fast for the classes