What changed when our cameras started remembering, with Hindsight A developer built Kaya, a construction-site safety copilot combining a Raspberry Pi 5 headset with YOLO11, YOLO-World, YOLO-Pose, Depth Anything V2, a PPE checker, a Docling-based RAG index, and Gemini Flash Lite, but found the pipeline had no memory beyond a 6- to 8-second frame buffer. After rejecting a Postgres event table as ill-suited to questions like "is the east scaffold still a problem?", the developer integrated Hindsight, an open source agent memory layer, to retain and recall hazard events with entity linking and semantic, keyword, graph and time-based retrieval. There's a line in our attention tracker that says, "If a hazard hasn't shown up in the frame for more than 5 seconds, mark it RESOLVED." It made sense when I wrote it. It also meant a worker could walk past an open floor edge at 9am and walk past it again at 2pm, and our system would act surprised both times. That line is the reason this post exists. What Kaya is Kaya is a safety copilot for construction sites. Workers wear a small headset we built: a Raspberry Pi 5 with an IMX219 camera, a mic, and a speaker. The site also has regular CCTV cameras. Every one of these streams goes to a single web server. On that server we run a pretty heavy stack: YOLO11 for people, vehicles, and obstacles, and YOLO-World for tools YOLO-Pose for 17 body keypoints, which we use for fall detection and head direction Depth Anything V2 for rough metric distance is that forklift 0.9 m away or 4 m? A PPE checker for hard hats, vests, gloves, and boots A Docling-based RAG index over the site safety manual and SOPs Gemini Flash Lite as the reasoning model, with Sarvam for speech in and out A worker can press a button and ask, "Is this ladder setup allowed?" and get a spoken answer in about 2 to 3 seconds that looks at the last few seconds of their camera feed and cites the SOP page it came from. All of this worked. It just had no memory. The problem with a system that only lives in the present Everything in the pipeline was built around "now." The CV side keeps a rolling 6- to 8-second ring buffer of frames. The chat side keeps a short list of recent turns in a Python list and trims it: Python max turns = self.settings. max history turns 2 if len self.conversation history max turns: self.conversation history = self.conversation history -max turns: And the attention tracker, which is the part I'm proudest of, throws hazards away once they leave the frame: python if age 5.0: entry "hazard" .state = HazardState.RESOLVED The attention tracker is the interesting bit. It doesn't just detect a hazard, it checks whether the worker has actually looked at it, by comparing their head yaw to the bearing of the hazard. If you look toward it for half a second, it counts as acknowledged. If you ignore it for 4 seconds, it escalates and the headset speaks up. So we had a system that knew whether you'd noticed something. And then five seconds later it forgot the thing existed. In practice this caused three problems: Repeat warnings. The same missing guardrail on the third floor got flagged fresh every time someone walked by. Workers started tuning it out, which is the worst outcome for a safety tool. No link between cameras. A CCTV camera would spot a hazard on one side of the site, and a worker's headset walking into that area had no idea. Each device was its own island. Useless answers to "has this happened before?" Site supervisors kept asking the chatbot things like "has anyone been near the crane without a helmet today?" The honest answer was "I have no idea, I live in an 8 second window." Why I didn't just write it to Postgres My first thought was to dump every hazard event into a table and query it. I actually started doing this. The problem is that the questions people ask aren't SQL-shaped. "Is the east scaffold still a problem?" needs the system to know that "east scaffold", "scaffold near gate 2" and zone scaffold b are the same place, that three separate sightings are one ongoing issue, and that yesterday's report matters more than one from three weeks ago. I would have ended up building my own entity linking, time weighting and search. That's a whole project on its own, and not the one I wanted to be working on. That's where Hindsight came in. It's an open source memory layer for agents. You retain things, you recall them, and it handles the part I didn't want to write: pulling out entities, linking facts over time, and searching with semantic, keyword, graph and time-based retrieval together. The Hindsight docs explain the retrieval side well if you want the details. One memory bank for the whole site The key decision was using one memory bank per site, not per device. Every headset and every CCTV camera writes into the same bank. That's the "hive mind" part. What one camera sees, every other device can know about. When the attention tracker escalates a hazard, or a fall is detected, or PPE fails, we retain it: python from hindsight client import Hindsight memory = Hindsight base url=settings.hindsight url SITE BANK = f"site-{settings.site id}" def remember hazard h: HazardAssessment, source: str - None: memory.retain bank id=SITE BANK, content= f"{source} saw {h.description} in zone {h.zone name}, " f"about {h.distance meters:.1f} m from the worker. " f"Severity {h.severity.name}. Acknowledged: {h.is acknowledged}." , source is something like helmet-07 or cctv-gate-2. We write plain sentences on purpose. Hindsight does the extraction, so I don't have to design a schema for every kind of event. On the question side, the pipeline already had a router that decides whether a question needs the safety manual. It's just a big list of regexes, and it runs in under a millisecond. I added one more branch: if the question is about history "today", "again", "before", "still", "has anyone" , we recall from the site bank before calling the vision model. if routing decision.needs memory: past = memory.recall bank id=SITE BANK, query=question memory context = m.text for m in past.results :5 Those memories go into the prompt next to the camera frames and the SOP chunks. The model now has three sources: what it sees right now, what the rules say, and what the site has seen before. Before and after Here's a real kind of exchange, before memory: Worker: Is this area safe to walk through? Kaya: I can see scaffolding to your left with no visible guardrail on the top level. Per the scaffolding SOP, guardrails are required above 1.8 m. Please avoid the area. That's fine. It's also exactly what it said the last four times. After adding Hindsight: Worker: Is this area safe to walk through? Kaya: The top level of this scaffold has had no guardrail since this morning. The gate camera flagged it at 9:40 and two other workers were warned here today. Avoid it and let your supervisor know it's still open. Same camera frames, same SOP. The difference is that the answer now knows this isn't new, and that makes it something a supervisor can act on. And the supervisor question that used to get a shrug now gets a list: which zones had PPE issues today, which cameras saw them, and whether they were acknowledged. We also changed the attention tracker a little. Before escalating, it recalls whether this worker has already acknowledged this hazard earlier. If they have, we skip the spoken alert and just show it on the dashboard. Fewer alerts, and the ones that do get spoken actually matter. Where this goes next The part I didn't expect when we started is how useful the memory bank is outside the safety use case. Every retained event is a short, timestamped description of something that happened on a real site, from a first-person or fixed camera, linked to the frames it came from. That's training data. We're building our own robots, and what they need most is a lot of real examples of how work actually happens on site and where things go wrong. The same shared memory that stops a worker from walking into an open edge is slowly becoming a record of how the site works, which we can use to teach machines to work there too. What didn't go well A few honest notes: I retained too much at first. Early on I retained every detection, and the bank filled up with junk like "person detected at 0.9 m" a thousand times. Now we only retain escalations, falls, PPE failures and anything a worker asks about. Being picky about what you save matters more than how you save it. Recall adds latency. It's not huge, but we're already at 2 to 3 seconds per answer and every extra call is felt. That's why memory only runs when the router decides the question needs it, not on every turn. Zone names have to be consistent. Hindsight is good at linking entities, but if one camera calls it scaffold b and another calls it east scaffold, you're making its job harder. We cleaned up our zone config so every device uses the same names. The regex router is a crutch. It's fast and predictable, which I like, but it will miss oddly worded questions. At some point it'll probably become a small classifier. What I'd tell someone building something similar Scope memory by where things happen, not by device. A per-device bank would have kept each camera on its own island. One bank for the site is what made it useful. Retain sentences, not rows. Let the memory layer do the extraction. You'll change your event types ten times, and plain text doesn't care. Only save what you'd want to be told about later. If you'd never ask about it, don't retain it. Decide when to recall. Treat memory lookups like RAG lookups and route them. Most questions don't need history. Use memory to stay quiet, not just to say more. Our biggest improvement wasn't smarter answers, it was fewer repeated alerts. If you're building agents that deal with the physical world, the "now" is only half the picture. Vision models are great at telling you what's in front of you. They're terrible at telling you if it was there yesterday. That's the gap agent memory fills, and for us it turned a smart camera into something closer to a coworker who has been on site all week. The code is here: https://github.com/shaurya-dogra/egocentric construction partner https://github.com/shaurya-dogra/egocentric construction partner