How we turned passive IP cameras into an intelligent alarm: a computer vision system that watches the live feeds, confirms when a real person appears in a protected area, and sends an annotated snapshot to the team in seconds.
The Challenge: Cameras That Record, But Never Warn
Most CCTV setups are passive. The cameras record around the clock, but nobody is watching the live feed - so an intruder is only discovered hours later, when someone scrubs back through the footage. The brief here was the opposite: know the instant a person enters a monitored area, with photographic proof, and without a constant stream of false alarms from passing cars, animals or swaying branches.
Our Solution: Computer Vision on the Live Feed
We built a system that connects straight to the existing IP cameras and turns a passive recording setup into an active, intelligent alarm. It pulls the live video, decides whether a human is really there, and pushes an annotated snapshot to the team within seconds - with no new hardware beyond a small cloud server.
How It Works (High-Level)
1. Live video ingest. The system pulls RTSP streams in real time from Dahua IP cameras connected to an NVR, straight into a dedicated processing server on Google Cloud.
2. Two-stage detection. A cheap motion filter runs first; only the frames that actually changed are sent to the neural network (more on why this matters below).
3. Person recognition. A YOLOv8 model locates and classifies people in the frame and attaches a confidence score to each detection (for example, 0.85).
4. Instant alert. The moment a person is confirmed inside a monitored zone, an annotated snapshot - bounding box, confidence, camera name and timestamp - is sent to the team over Telegram.
The Technology Behind It: Two-Stage Detection
Running a neural network on every frame of every camera, 24/7, would be wasteful and expensive. So we filter first:
- Stage 1 - Motion filter (MOG2). A lightweight background-subtraction algorithm (OpenCV MOG2) looks for movement. Static frames - an empty yard, a locked gate - are discarded instantly, before any AI runs. This step is computationally cheap and removes the vast majority of frames.
- Stage 2 - Person detection (YOLOv8). Only frames with real motion reach the neural network. A YOLOv8 model (Ultralytics, trained on the COCO dataset) then classifies and locates people, drawing a bounding box and attaching a confidence score.
The result: the accuracy of a modern detection model at a fraction of the compute cost.
Smart Zones: Reacting Only Where It Matters
A camera pointed at a yard also sees the street behind it. To avoid constant false alarms, every camera has:
- Regions of interest (ROI). Configurable polygon zones mean the system only reacts to activity inside the areas you care about, and ignores the pavement, the road, or a neighbour's driveway.
- Exit zones. If a person is detected leaving through a known exit, alerts are briefly suspended, so staff heading home don't trip the alarm.
No Alert Floods
A detection system that cries wolf gets muted within a day. To keep every alert meaningful, there are two layers of throttling:
- Per-camera cooldown stops a single camera from firing repeatedly for one continuous event.
- Global cooldown caps the overall alert rate, so a busy moment never buries the team in notifications.
Real-Time Alerts, With Proof
When a person is confirmed, the team gets a Telegram message within seconds containing the snapshot with the detected person boxed, the confidence score, and the camera name with a timestamp. No logging into a separate app, no scrubbing footage - the evidence lands on the phone that is already in their pocket.
The Architecture (For the Curious)
- Python + OpenCV for video processing and the motion stage.
- YOLOv8 (Ultralytics) for person detection.
- FastAPI exposes a REST API for live status and configuration.
- Multi-threaded by design - each camera runs in its own thread, feeding a producer-consumer queue, so a slow or dropped stream never blocks the others.
- Google Cloud Platform - the whole service runs as a systemd service on a GCP virtual machine, so it restarts itself and survives reboots.
- Tailscale VPN for secure remote access to the server and cameras, with no ports exposed to the public internet.
Why This Matters
Security hardware you already own, made intelligent with software. Instead of a guard watching a wall of monitors - or nobody watching at all - a computer vision pipeline watches every feed, every second, and only interrupts a person when there is a real human in a place they should not be. Fewer false alarms, faster response, and a full audit trail of annotated snapshots.
Could Computer Vision Work for Your Operation?
Whether it is security monitoring, counting, quality inspection, or spotting events a person would miss - if a camera can see it, there is a good chance we can build software to act on it automatically.
About Inigra Software House. We are a European software house specializing in AI integration, computer vision, and custom software development. We help businesses turn the data and hardware they already have into tools that work for them.


