Nidal Siddique Oritro
← Back to posts

Yolo26 Onnx For Frigate - Tested and Implemented

homelab 12 min read
Yolo26 Onnx For Frigate - Tested and Implemented
Summary YOLO26 doesn't work in Frigate out of the box. Here is why, how i exported it so it does, and how all 10 models did against YOLOv9 on a real recording.

YOLO26 is the newest model from Ultralytics, released January 2026, and it doesn’t work in Frigate out of the box.

I have been running YOLOv9 on Frigate since the 0.16 upgrade ( the export post is here ), and last night i finally finished exporting every v9 size, at both 320 and 640. So the obvious next question was, can i run YOLO26 instead, and is it actually better. I went down the rabbit hole, exported all of them, tested all of them in Frigate, and this post is the whole thing.

My setup, for context:

text
- Frigate 0.17.1, tensorrt image
- ONNX detector on CUDA, model_type yolo-generic
- Nvidia RTX 2080 Ti
- Current model: YOLOv9-S, 640

The Problem With YOLO26

There are two problems, one political and one technical.

The political one first. Frigate doesn’t officially support Ultralytics models. Not YOLO26, not YOLO11, not YOLOv8. A maintainer said it plainly on the YOLO26 feature request, “Frigate is not able to officially support Ultralytics models at this time”, and pointed to PR #10717 for the reasons. The detector docs don’t mention YOLO26 at all. So whatever you do here, you are on your own if a Frigate upgrade breaks it.

The technical one is the actual blocker. YOLO26 ships with two detection heads:

HeadOutput shapeNeeds NMSDefault export
One-to-one ( end-to-end )[1, 300, 6]NoYes
One-to-many ( classic )[1, 84, 8400] at 640YesNo

The whole selling point of YOLO26 is the end-to-end head, it does its own filtering and spits out up to 300 final boxes, no NMS needed. That is also the default when you export it. But Frigate’s yolo-generic post-processing expects the classic YOLO layout, 84 rows ( 4 box coordinates + 80 COCO class scores ) by 8400 candidates, and then runs its own NMS on top. Feed it [1, 300, 6] and it reads the wrong numbers as the wrong things.

What People Ran Into

I went through the Frigate GitHub discussions before touching anything, and the reports are all over the place:

  • Confidence scores of 6000% and 10000%. Multiple people reported this, along with “random bicycles everywhere” and “a bike in the sky”. That is exactly what happens when Frigate reads a [1, 300, 6] tensor as if it were [1, 84, 8400], box coordinates end up getting treated as class scores.
  • No detections at all. One user exported YOLO26-M as both ONNX and OpenVINO and got nothing back from Frigate.
  • Double the latency. The same user saw YOLO26-M at ~40 ms against YOLO11-M at ~20 ms on an Intel Iris Xe iGPU.
  • Monkey patching. One person patched the PyTorch code before export to force the old output format, and published prebuilt models for it.
  • The actual fix, buried in a reply. Someone eventually pointed out that you don’t need any patching, export it with end2end=False and it works, and on their hardware it was faster than v9 and v11.

So the information was there, just spread across three threads, with no tested set of models at both sizes, and no numbers on real footage. That’s the gap i wanted to fill.

How I Implemented It

Same approach as my v9 repo, a small container that does the export on CPU and drops the .onnx file out the other side. Unlike v9, there is nothing to patch in the source, Ultralytics does the heavy lifting. The only part that matters is end2end=False.

text
FROM python:3.11 AS build
RUN apt-get update && apt-get install --no-install-recommends -y libgl1 libglib2.0-0 && rm -rf /var/lib/apt/lists/*
COPY --from=ghcr.io/astral-sh/uv:0.8.0 /uv /bin/
WORKDIR /work
RUN uv pip install --system --index-url https://download.pytorch.org/whl/cpu torch==2.8.0 torchvision==0.23.0
RUN uv pip install --system ultralytics==8.4.175 onnx onnxslim onnxruntime --extra-index-url https://download.pytorch.org/whl/cpu --index-strategy unsafe-best-match
ARG MODEL_SIZE
ARG IMG_SIZE
RUN python3 -c "from ultralytics import YOLO; YOLO('yolo26${MODEL_SIZE}.pt').export(format='onnx', imgsz=${IMG_SIZE}, end2end=False, simplify=True, dynamic=False, batch=1)"
FROM scratch
ARG MODEL_SIZE
ARG IMG_SIZE
COPY --from=build /work/yolo26${MODEL_SIZE}.onnx /yolo26-${MODEL_SIZE}-${IMG_SIZE}.onnx

Run it with --build-arg MODEL_SIZE=s --build-arg IMG_SIZE=640 ( sizes are n s m l x ) and --output ., same as the v9 one.

end2end=False switches the export to the one-to-many head, and the output comes back in the classic layout. I checked the output shape of every single file before calling it exported:

SizeOutput shape
640[1, 84, 8400]
320[1, 84, 2100]

All 10 builds ( 5 sizes × 2 resolutions ) took about 6 minutes in total on CPU, after the first build cached the torch layers. Each one was 20 to 40 seconds.

Frigate config is the same yolo-generic block you would use for v9, just point it at the new file:

yaml
model:
  path: /config/models/yolo26-s-640.onnx
  model_type: yolo-generic
  input_tensor: nchw
  input_pixel_format: rgb
  input_dtype: float
  width: 640
  height: 640

For 320, change the path and both width and height to 320. Don’t forget the second one, the size in the config has to match the size in the file name.

How I Tested It

I didn’t want to swap models on my live Frigate ten times, that’s ten restarts and ten gaps in recording. So the testing went in three stages, each one closer to the real thing.

Stage 1, bench check. I ran every model inside the Frigate container using Frigate’s own onnxruntime ( CUDA ) and Frigate’s own post_process_yolo function, which is the exact code yolo-generic uses. Two images, the standard Ultralytics bus photo and one live camera frame. The point here was just to see if the scores came back sane.

Stage 2, replay on a real recording. I picked a 70 second recording from one of my indoor cameras, evening, three people in the frame the whole time. Dining room, dim light, people partially hidden behind each other and the furniture, which is a hard scene and honestly a good one for this. I sampled 421 frames out of it and ran all 10 YOLO26 models plus YOLOv9-S and YOLOv9-C over the same frames. A person counts when the score is 0.5 or above.

Stage 3, actual Frigate. A throwaway Frigate instance, same 0.17.1-tensorrt image, same ONNX on CUDA, with that same recording looped as its only camera. Detect at 1280x720, 5 fps, recording off. Each model got loaded on its own, ran for 150 to 240 seconds, then i pulled Frigate’s own stats and events from its API. My live Frigate was never touched.

The Results

Bench check

Every model came back clean. All scores between 0 and 1, the right classes ( 4 people and a bus/car on the bus photo ), no phantom bicycles anywhere. The raw ONNX latency on the 2080 Ti:

Model320640
YOLO26-N3.0 ms4.4 ms
YOLO26-S3.3 ms6.7 ms
YOLO26-M6.1 ms16.9 ms
YOLO26-L8.7 ms15.3 ms
YOLO26-X10.4 ms28.0 ms
YOLOv9-S—9.2 ms

These are ONNX on the CUDA provider, not TensorRT, so don’t compare them with the Ultralytics T4 numbers.

In Frigate

Every model loaded, every model created person events, no score over 1, zero errors in the logs.

ModelSizeRunFrigate inferencePerson eventsAvg top scoreMax top score
YOLO26-N320150 s4.1 ms60.840.90
YOLO26-N640150 s13.9 ms80.810.89
YOLO26-S320240 s5.3 ms90.860.93
YOLO26-S640240 s14.3 ms130.840.91
YOLO26-M320150 s7.8 ms60.840.94
YOLO26-M640150 s16.3 ms70.850.94
YOLO26-L320150 s9.7 ms50.870.94
YOLO26-L640150 s24.0 ms60.860.94
YOLO26-X320150 s17.0 ms40.920.94
YOLO26-X640150 s28.1 ms40.920.95
YOLOv9-S640240 s16.3 ms100.900.93

Couple of things to read this right.

  • Frigate’s inference number is not the raw model speed. It includes Frigate’s own overhead around the model, which is why N, S and M at 640 all sit around 14 to 16 ms even though the raw model times are very different. At 640 that overhead floor is most of the number. At 320 the floor drops a lot, which is why 320 looks so much faster in Frigate.
  • Don’t compare event counts across rows with different run times. The clip loops, so a longer run means more events. Only compare rows with the same Run value.

On the recording

This is the table that actually matters, because it’s how many of the three people each model could find, frame by frame.

ModelSizeFrames with ≥1 personAvg people found per frame ( out of 3 )Frames with all 3Avg person scoreRaw latency
YOLO26-N32013.5%0.190%0.612.4 ms
YOLO26-N64011.9%0.190%0.654.6 ms
YOLO26-S32038.2%0.670%0.723.3 ms
YOLO26-S64057.7%1.011.0%0.756.7 ms
YOLO26-M32071.7%1.260.2%0.765.1 ms
YOLO26-M64062.0%1.130.7%0.7812.0 ms
YOLO26-L32071.3%1.341.9%0.776.6 ms
YOLO26-L64065.8%1.231.9%0.7714.6 ms
YOLO26-X32076.2%1.443.1%0.7610.4 ms
YOLO26-X64067.9%1.293.3%0.7828.3 ms
YOLOv9-S64039.7%0.710.2%0.788.5 ms
YOLOv9-C64048.5%0.840.2%0.7715.9 ms

What i read from it:

  • Nobody finds all three people. The best model, X, gets all three in about 3% of frames. It’s a dim, crowded indoor scene with people overlapping, so this is a hard test, not a typical doorbell camera.
  • N is not usable on a scene like this. It finds a person in roughly 1 out of 8 frames.
  • 320 is as good as 640 here, sometimes better. M, L and X all found more people at 320 than at 640. I think it’s because the people take up a big part of the frame in this camera, so shrinking the image doesn’t lose them, but i haven’t proved that. On a wide outdoor camera with people far away, i would expect 640 to win.
  • The scores look similar across the board. Avg person score sits around 0.75 to 0.78 for everything from S up, v9 included. The difference is in how many people get found, not how sure the model is about the ones it finds.

How It Compares To YOLOv9

First the official numbers, COCO mAP 50-95 at 640, from Ultralytics and the YOLOv9 repo:

TierYOLOv9YOLO26
Tiny / NanoT: 38.3N: 40.9
SmallS: 46.8S: 48.6
MediumM: 51.4M: 53.1
LargeC: 53.0L: 55.0
ExtraE: 55.6X: 57.5

On paper YOLO26 is about 1.7 to 2.6 mAP ahead at every size. Not huge.

On my recording, the gap was a lot bigger than the paper says:

YOLOv9-S ( what i run )YOLOv9-CYOLO26-S 640YOLO26-M 320
Frames with ≥1 person39.7%48.5%57.7%71.7%
Avg people found0.710.841.011.26
Frigate inference16.3 ms—14.3 ms7.8 ms

YOLO26-S at the same size as my current model finds people in 18 more frames out of every 100, and it’s slightly faster in Frigate. Even the bigger YOLOv9-C, which is roughly four times the compute of v9-S, loses to YOLO26-S. And YOLO26-M at 320 finds people in almost twice as many frames as my current v9-S, at half the inference time.

The things v9 still has going for it:

  • Official support. v9 is what the Frigate team tests against, with an export guide in the docs. YOLO26 works because of an export flag, and if a future Frigate release changes how yolo-generic parses output, that’s on me to fix.
  • License. v9 is GPL-3.0, YOLO26 is AGPL-3.0. For a homelab it doesn’t matter. If you are building something you ship to other people, read it.

Also, to be clear about what this is, it’s one 70 second clip from one indoor camera, at night. It’s a real test on real footage, but it’s not a benchmark. I’m gonna run it on my outdoor camera during the day before switching my live Frigate over, and if you have a different kind of scene, test it on your own footage first.

Download

Everything is on GitHub, both 320 and 640, all five sizes, each one exported, shape checked, and tested in Frigate:

YOLO26-onnx on GitHub

  • N, S, M and L are in the repo itself.
  • X is over GitHub’s 100 MB file limit, so it’s a release download, X 640 and X 320. SHA-256 for each is in the release notes.

And if you are still on v9, all the v9 models at both sizes are here: YOLOv9-onnx on GitHub.

If you only want my pick, start with YOLO26-M at 320. It found the most people per millisecond in my test, and it’s faster than the v9-S most people are running right now.

Published October 12, 2026

Nidal Siddique Oritro

Oritro Ahmed

Software Engineer turned Engineering Manager. Writing about software, teams, homelab, and AI.

Comments