Computer Vision

Computer vision for construction-site safety in the UAE: PPE and restricted-zone monitoring

How camera-based PPE detection and restricted-zone alerts work in practice, what they cost to run, and the questions to ask before putting a model on a live site in Dubai's heat and dust.

Most safety incidents on a construction site are not mysterious. A worker without a helmet in a lifting zone. Someone stepping past a barrier to save thirty seconds. A supervisor cannot watch forty people across six floors, but a camera can, and an object-detection model can watch the camera. This article explains how the two systems we build most often, PPE detection and restricted-zone monitoring, actually work, and what we have learned deploying them on UAE sites.

PPE detection: what the model is really doing

A PPE detector is an object-detection model that has been taught the classes that matter on site: person, helmet, high-visibility vest, gloves, sometimes safety glasses and harnesses. For each camera frame it returns a box and a confidence score for every instance it sees. The safety logic then runs on top of those boxes: for every person box, is there a helmet box overlapping the top of it? Is there a vest box overlapping the torso? If not, that person is non-compliant, and the frame, the timestamp and the camera ID go into an event log.

Two design decisions matter more than the model's headline accuracy.

Open-vocabulary vs fixed classes

A classic YOLO model knows only the classes it was trained on. Add a new PPE requirement (say, a specific colour of vest for a subcontractor) and you retrain. An open-vocabulary detector such as YOLO-World takes a text description of the class at runtime instead. Our PPE system uses this approach so a safety officer can add "orange vest" or "welding mask" from the web console, upload a few reference images, and have it monitored the same afternoon. The trade-off is speed and a small drop in precision on unusual angles, which is why we still fine-tune a fixed-class model for sites with stable requirements and dozens of cameras.

Where the compliance rule lives

Keep the rule ("helmet required in zones A and B, harness required above level 3") outside the model, in configuration. Rules change per site, per phase and per contractor. Models should not.

Restricted zones: geometry, not magic

A restricted-zone monitor is simpler than it sounds. A person detector (we use YOLOv11 in its CPU-optimised form, so it runs without a GPU) finds every person in the frame. The user has drawn one or more polygons on the camera image: the crane swing radius, an open shaft, the area behind the concrete pump. For each detected person, the system takes a reference point, usually the bottom-centre of the box where the feet are, and tests whether it falls inside a polygon. If it does, that is an intrusion.

The details that make it usable on a real site:

  • Cooldown. A person standing inside a zone should generate one alert, not thirty per second. We use a short cooldown per zone, and track the person so a re-entry is a new event.
  • Evidence. Every alert stores the frame with the box and the zone overlay drawn on it. Safety managers need to show the image in a toolbox talk, not describe it.
  • Status. Alerts are active until someone resolves them from the dashboard, so nothing gets lost in a stream of notifications.
  • Reconnection. Site RTSP cameras drop. The stream reader has to reconnect on its own, and the dashboard has to show which cameras are currently live.

The UAE-specific problems

Models trained on European or North American footage meet three things on a Gulf site they have not seen much of.

  1. Glare and heat haze. Midday sun on light-coloured surfaces washes out contrast; afternoon haze softens edges. Confidence thresholds tuned at 9am will miss detections at 1pm. We calibrate per camera using footage from the full day, and we run a lower threshold with a temporal check (the same person seen without a helmet across several consecutive frames) rather than a single high-confidence frame.
  2. Dust on the lens. A camera that was sharp on installation day is soft two weeks later. The system should track its own detection rate per camera and flag a drop, because a silent camera looks exactly like a compliant site.
  3. Clothing. Long sleeves, head coverings under helmets, and dust masks are the norm. A model that learned "helmet" as "hard hat directly on hair" will underperform. This is a fine-tuning problem, and it is why we always ask for a few hours of the client's own footage before promising any accuracy number.

What it costs to run

People expect the model to be the expensive part. It usually isn't. A CPU-only person detector handles a handful of cameras on one modest machine; a GPU box handles dozens. The costs that actually decide the project are camera coverage (can the existing CCTV see the zones you care about at a usable angle?), network (can the video reach the machine reliably?), and process (who gets the alert, and what are they obliged to do with it?). If the third one has no answer, the system becomes a log nobody reads.

Questions to ask any vendor, including us

  • Can you show me detections on my footage before we sign, not on a demo video?
  • What happens when a camera goes offline, and how will I know?
  • Can I add a new PPE item or redraw a zone without calling you?
  • Where is the video processed, and does any of it leave the site?
  • What is the false-alarm rate at 1pm in July, not the accuracy figure from the paper?

Good answers to those five questions matter more than any benchmark. If you want to see how we answer them, the PPE detection and restricted-zone monitor pages walk through the architecture, and we are happy to run a scoped trial on your cameras.

Running a site that needs this?

We build PPE and restricted-zone monitoring on your existing cameras, and we scope it on your footage before you commit.

Start a project →