Computer Vision
YOLO vs open-vocabulary detection: how to choose for a real deployment
We have shipped both fixed-class YOLO models and open-vocabulary detectors like YOLO-World and OWL-ViT. Here is the decision we actually make on each project, with the trade-offs that matter outside a benchmark.
Five years ago the answer to "which object detector?" was always some version of YOLO. Today there is a second family, open-vocabulary detectors such as YOLO-World, OWL-ViT and Grounding DINO, that find objects from a text description without any training. We use both. This post is the decision we actually make at the start of a project, written down.
The two families in one paragraph each
Fixed-class detectors (YOLOv8, YOLOv11 and friends) are trained on a labelled dataset with a closed list of classes. They are very fast, run comfortably on a CPU in their small variants, and, once fine-tuned on your footage, are very precise. Their limitation is that the class list is frozen at training time. New class, new dataset, new training run.
Open-vocabulary detectors (YOLO-World, OWL-ViT, Grounding DINO) pair a vision backbone with a text encoder. You give them a prompt ("AED banknote", "orange safety vest", "open cash drawer") and they find matching regions with no training at all. They are slower, usually want a GPU, and their precision on unusual objects or angles is lower than a fine-tuned model. Their strength is flexibility: the class list is a string you can change at runtime.
Three projects, three different answers
Restricted-zone monitor: fixed-class YOLOv11
The only class that matters is person, which every pretrained YOLO already knows well. The site had no GPU, several cameras, and needed a live 30 FPS display. A CPU-optimised YOLOv11 nano model was the obvious fit. More on that build.
PPE detection: open-vocabulary YOLO-World
PPE requirements change per site, per contractor and per phase. Retraining every time a client added "welding mask" was not sustainable. YOLO-World let us expose an "add PPE item" button in the console: type the name, upload a few reference images, done. We accepted a small precision cost for a large operational win. More on that build.
Currency detector: OWL-ViT plus OCR
The job was to notice banknotes left on a counter and tell AED from USD from EUR. Banknotes are not in any standard detection dataset, and collecting a labelled set of every denomination at every angle was more work than the project justified. OWL-ViT found notes from the prompt "banknote" reliably, and OCR on the crop handled the currency classification. The crops it collected became labelled YOLOv8 datasets for a fixed-class model, which is a common pattern: start open-vocabulary, harvest data, graduate to fixed-class. More on that build.
The decision, as a list of questions
- Is the object a common class? Person, car, truck, bicycle, dog: a pretrained fixed-class model already handles it. Stop here and use YOLO.
- Will the class list change after launch? If the client will add categories themselves, open-vocabulary saves you a retraining pipeline and them a support ticket.
- Do you have labelled data, or can you get it cheaply? A few hundred labelled frames of the client's own footage turns a fine-tuned YOLO into the most accurate option by a wide margin. No data means open-vocabulary, at least to start.
- What hardware is on site? CPU only, or an edge box, points to a small YOLO. A GPU server opens up the open-vocabulary models.
- How expensive is a false alarm? If every alert wakes a security guard, precision wins and you should fine-tune. If alerts feed a dashboard someone reviews in the morning, recall and flexibility win.
Things benchmarks don't tell you
- Prompt sensitivity. Open-vocabulary detectors can behave differently for "helmet" vs "hard hat" vs "safety helmet". Budget time to test prompts on real footage, and store the winning prompts in configuration.
- Small objects. Both families struggle with objects that occupy a few dozen pixels. Camera placement fixes more of this than model choice does.
- Temporal logic beats per-frame accuracy. Requiring the same detection across several consecutive frames removes most false alarms at almost no cost. We do this on every project regardless of the detector.
- Versioning. Whichever you choose, pin the model weights and record which version produced each alert. When someone asks why the system flagged a frame six weeks ago, you need to be able to reproduce it.
Our default
If we had to pick one rule: start open-vocabulary when the classes are unusual or unstable, collect the crops it finds, and fine-tune a small YOLO once you have data and the requirements have settled. You get the flexibility early and the precision and speed later, and the migration is a configuration change rather than a rewrite.