Every security team has the same quiet frustration: an alert fires, someone spends 40 minutes diagnosing it, fixes it, and moves on. Three weeks later, the exact same type of alert fires again — and the team starts from zero, because nothing remembers what happened last time.
That's the gap we set out to close with SentinelMind, our project for HackwithHyderabad 3.0.
The problem with most AI security tools
Most "AI-powered" triage tools are really just a chatbot wrapped around an alert feed. They reason well in the moment, but they have no memory — every incident is treated as if it's the first one they've ever seen. That means teams keep re-solving solved problems, and any institutional knowledge lives in someone's head instead of the system.
What we built
SentinelMind is an incident response agent that uses Hindsight, a persistent memory layer, to remember every incident it's triaged: what type it was, what fix was applied, whether it worked, and how long it took. When a new incident comes in, the agent doesn't just reason from scratch — it checks memory first.
If it finds a similar past incident, it tells you directly: "This matches an incident from three weeks ago — the fix that worked was X, and it took 15 minutes." If it's never seen anything like it, it says so honestly instead of pretending to be confident. Proving memory actually helps
The part we're proudest of isn't the triage logic — it's the Memory Impact dashboard, which tracks resolution time across every incident the agent has handled. Early incidents (before enough memory built up) take 40+ minutes to resolve. Later incidents, where the agent finds a strong memory match, resolve in under 15. Watching that curve trend downward across a live demo is a far more convincing proof point than any explanation could be.
The hardest technical problem wasn't what we expected
We assumed wiring up the LLM and the memory layer would be the hard part. It wasn't. The actual challenge was stopping the model from inventing statistics — success rates, average resolution times — when it didn't have real numbers to draw on. Our fix: every number shown in the app is computed directly from stored incident records in code. The LLM's only job is to explain those numbers in plain language, never to generate them itself. That one architectural decision made the whole system trustworthy instead of just plausible-sounding.
Stack: Node.js/Express backend, React frontend, Groq (running openai/gpt-oss-120b) for reasoning, Hindsight for persistent memory, Docker for packaging.
What's next: the current version handles six incident types with seeded historical data. The natural next step is wiring it to a real alert source (a SIEM feed or parsed network capture data) instead of manually submitted incidents — same architecture, real data.
If you're building something in the "AI agent with memory" space for a hackathon, happy to talk shop about the Hindsight integration or the memory-matching logic — feel free to reach out.