{"slug": "argus-brightsign-machine-vision-it-s-not-difficult", "title": "Argus, BrightSign machine vision - it's not difficult!", "summary": "BrightSign released Argus, its flagship audience-measurement extension for edge AI on its LS5, XT5, and XD6 players, which runs YOLOX and RetinaFace neural networks in parallel on the NPU to detect and track people and gaze in under 15ms per model at 3.5 watts. The open-source reference app streams JSON analytics over MQTT and Prometheus endpoints without sending images off-device, and BrightSign's developer advocate highlights the companion HTML5 demo that visualizes detection in real time.", "body_md": "I’ve written before about [what a BrightSign extension actually is](https://blog.herlein.com/post/developing-brightsign-extension/) and about [why edge AI has to run locally on signage](https://blog.herlein.com/post/playback-to-perception/). This post is the “show, don’t tell” follow-up. If you want to see what all that theory looks like as a running, watchable thing, there’s exactly one place to start: [ Argus](https://github.com/brightsign/argus-audience-measurement-extension).\n\nArgus is the BrightSign flagship reference application for audience measurement on the NPU, and it’s genuinely the best example I can point you at of “here’s what an edge AI player can actually do.” But it didn’t arrive out of nowhere. Behind it is a trail of smaller repos - each one a single building block, a single lesson learned, a single “can we even do this?” question answered before we bet on the bigger thing. I want to walk you through both: how to actually stand up Argus and watch it work, and then the history of everything that led to it. Because that history is useful. It’s the map of how you’d build your own.\n\n## Start Here: Argus\n\nArgus watches a camera feed and, entirely on-device, tells you how many people are in front of your screen, whether they’re actually looking at it, how long they dwell, when they enter and exit, and which direction they’re moving. Two neural networks - YOLOX for person detection, RetinaFace for face detection - running in parallel on the NPU, fused with ByteTrack so it can follow individual people across frames. Sub-15ms inference per model, all in about 3.5 watts.\n\nArgus Panoptes - the hundred-eyed watcher. Fitting name for a computer-vision system.\n\nGetting it running is refreshingly boring, in the best way:\n\n- Grab the\n[latest Argus BSFW release](https://github.com/brightsign/argus-audience-measurement-extension/releases/latest)and drop it on the root of an SD card. - Plug in a USB webcam (or point it at an RTSP stream if you want to go fancy).\n- Boot the player.\n- All on a BrightSign LS5, XT5, or XD6 player.\n\nThat’s it. From another machine on the network:\n\n```\nmosquitto_sub -h <PLAYER_IP> -t 'bs/argus/#' -v\n```\n\nAnd you’ll start seeing JSON like this stream by, live, every frame:\n\n```\n{\"schema\":\"analytics/v7.0\",\"people\":2,\"gaze\":1,\"fps\":29,\"tracks\":[...]}\n```\n\nNo cloud round-trip. No images leaving the box - it does face *detection* to compute gaze, not face *recognition*. Nobody’s identity gets stored anywhere. Just counts, durations, and attention rates coming out the other side over MQTT and a Prometheus `/metrics`\n\nendpoint. I harped on this in the [playback-to-perception post](https://blog.herlein.com/post/playback-to-perception/) and I’ll keep harping on it: that architecture choice is the whole ballgame if you ever want to defend this to a privacy-conscious customer, and you will.\n\n### Actually Watching It Think\n\nRaw MQTT messages are great for engineers, useless for convincing a skeptical stakeholder in a conference room. That’s exactly what [ argus-attention-demo-html](https://github.com/brightsign/argus-attention-demo-html) is for. It’s a small HTML5/JS app that runs full-motion video on the player alongside a picture-in-picture window showing exactly what the camera AI “sees” - bounding boxes drawn live around every face, green if that person is looking at the screen, red if they’re not.\n\nStanding it up is a `make prep && make build && make publish`\n\n, which drops everything you need into an `sd/`\n\nfolder - copy that to a card, install the Argus BSMP alongside it, boot, and watch. This is the demo I’d put in front of anyone who says “edge AI in signage” is marketing vaporware. Point a webcam at yourself, look away, look back, and watch the box change color in real time. It’s hard to argue with your own face turning green and red on a screen.\n\n## How We Got Here: The Repo Family\n\nArgus didn’t get built in one shot. It’s the sum of a bunch of smaller, single-purpose repos, each one proving out a piece before we combined them. If you’re the type who learns by taking things apart - I am - this list is worth going through in order, because it’s basically a guided tour of NPU development on BrightSign hardware, from “what even is this chip” to “full production reference app.”\n\n- The primer. What an NPU actually is, why it matters for edge inference, and a map of every other repo in this family. If you’re starting from zero, start here.[brightsign-npu-general](https://github.com/brightsign/brightsign-npu-general)- The first single-model building block: RetinaFace face detection plus the “is this person looking at the screen?” logic. This is the gaze half of Argus, in isolation.[brightsign-npu-gaze-extension](https://github.com/brightsign/brightsign-npu-gaze-extension)and[simple-gaze-detection-html](https://github.com/brightsign/simple-gaze-detection-html)- Two thin demo wrappers around the gaze extension, one built as an HTML5 app and one as a BrightAuthor:connected presentation, so you can see the gaze BSMP work regardless of which authoring path you use.[simple-gaze-detection-presentation](https://github.com/brightsign/simple-gaze-detection-presentation)- Object detection with selectable classes and confidence thresholds. Point it at a shelf, a doorway, a queue - whatever object matters to your use case.[brightsign-npu-object-extension](https://github.com/brightsign/brightsign-npu-object-extension)- The BrightAuthor:connected demo that shows the object extension changing playback the instant it detects the object you told it to watch for.[simple-object-detection-presentation](https://github.com/brightsign/simple-object-detection-presentation)- This one I still think is underrated: gaze-triggered speech-to-text using a Whisper encoder-decoder model, running on-device. Someone looks at the screen, the mic starts listening, and you get a transcript. No cloud, no wake-word service, no third-party voice API.[brightsign-npu-voice-extension](https://github.com/brightsign/brightsign-npu-voice-extension)- The HTML demo for the voice extension, so you can watch that gaze-triggers-listening pipeline work end to end.[simple-voice-detection-html](https://github.com/brightsign/simple-voice-detection-html)- The unglamorous one that made every other repo on this list easier to build. It’s a tiny web server that shows you the annotated frame the model is currently looking at, in a browser, so you’re not debugging computer vision blind over a serial console. Every one of the extensions above got easier to develop once this existed.[bs-image-stream-server](https://github.com/brightsign/bs-image-stream-server)\n\nNotice the pattern. Each repo answers one question - can we detect gaze, can we detect objects, can we trigger a mic off attention, can we actually *see* what the model sees while we’re building it - before Argus stitched person detection, gaze, tracking, and dual-output telemetry into one coherent production-grade application. That’s not an accident. That’s how you build anything real: small, provable pieces first, then integrate.\n\n## The Payoff: It’s Not Hard to Build These\n\nHere’s the thing I actually want you to take away from all of this. Every repo on that list, including Argus itself, is *just an extension*. Same mechanism I described in my [last extension post](https://blog.herlein.com/post/developing-brightsign-extension/) - a squashfs filesystem, a `bsext_init`\n\nscript with `start`\n\n/`stop`\n\n/`run`\n\n, old-school SysV init. Nothing exotic. If you can write a Linux daemon, you can write a BrightSign extension.\n\nDon’t take my word for it - go do it yourself. [ bs-workshop-extension](https://github.com/BrightDevelopers/bs-workshop-extension) is a self-guided, ~3.5-hour, hands-on workshop that walks you through building a trivial “Hello BrightSign” Java extension - HTTP server, JSON uptime response - from a bare unsecured player all the way through build, package, deploy, verify, and iterate. It even ships a companion HTML app,\n\n[, so you see the extension’s output rendered on the actual screen, not just in a terminal. The workshop is deliberately trivial on purpose - the point isn’t the app, it’s that you walk the](https://github.com/BrightDevelopers/bs-extension-workshop-html-app)\n\n**bs-extension-workshop-html-app*** entire*development loop yourself, once, so it stops being mysterious.\n\nAnd that loop - build, package, deploy, curse, iterate - is exactly the same loop that produced Argus. There is no secret extra step that only BrightSign engineers know about. I promise you.\n\n## And the NPU Really Isn’t That Complicated Either\n\nThe other myth I want to put down: that NPU development requires some rarefied deep-learning wizardry you need a PhD for. It doesn’t. Go read [brightsign-npu-general](https://github.com/brightsign/brightsign-npu-general) again with fresh eyes - it’s a page and a half of “here’s what an NPU is, here’s why it’s fast and low-power, here’s the model zoo you pull from.” The heavy lifting - the actual neural network architectures, YOLOX, RetinaFace, Whisper - is open, published, and well-documented upstream. Your job as the developer is wiring: get frames to the model, get results out, do something useful with them. That’s the same job you’d have wiring up any other sensor. It’s not magic. It just *looks* like magic because most people have never seen the inside of the box.\n\nThink about that for a minute. A chip that draws under four watts, sitting idle in millions of already-deployed signage players, capable of running person detection, face detection, and speech-to-text simultaneously - and the on-ramp to using it is a workshop you can finish in an afternoon.\n\n## Conclusion\n\nArgus is NOT a finished, production-grade thing. It’s a working demo that shows the power of the NPU and edge computing. Watch it work with the attention demo, subscribe to the MQTT feed, and you’ll get it in about five minutes. But if you want to actually *build* something like it, don’t stare at Argus and feel intimidated. Walk the trail that got us there: `brightsign-npu-general`\n\nfor the concepts, the single-model extensions for the individual building blocks, `bs-image-stream-server`\n\nso you’re not debugging blind, and `bs-workshop-extension`\n\nto prove to yourself the whole development loop isn’t scary. Every one of those repos was somebody’s “can I even do this” experiment before it became a reference architecture.\n\nThe NPU is sitting there, idle, in hardware that’s already deployed. Go point it at something. And if you build something cool with it, [drop me a note on LinkedIn](https://www.linkedin.com/in/gherlein/) - I genuinely want to see it!", "url": "https://wpnews.pro/news/argus-brightsign-machine-vision-it-s-not-difficult", "canonical_source": "https://blog.herlein.com/post/argus-and-the-npu-repo-family/", "published_at": "2026-08-25 08:00:01+00:00", "updated_at": "2026-08-25 21:14:41.395320+00:00", "lang": "en", "topics": ["computer-vision", "artificial-intelligence"], "entities": ["BrightSign", "Argus", "YOLOX", "RetinaFace", "ByteTrack", "LS5", "XT5", "XD6"], "alternates": {"html": "https://wpnews.pro/news/argus-brightsign-machine-vision-it-s-not-difficult", "markdown": "https://wpnews.pro/news/argus-brightsign-machine-vision-it-s-not-difficult.md", "text": "https://wpnews.pro/news/argus-brightsign-machine-vision-it-s-not-difficult.txt", "jsonld": "https://wpnews.pro/news/argus-brightsign-machine-vision-it-s-not-difficult.jsonld"}}