{"slug": "i-built-an-l1-ticket-deflector-that-knows-when-to-stop-langgraph-human-in-the", "title": "I built an L1 ticket deflector that knows when to stop: LangGraph + human-in-the-loop", "summary": "A developer open-sourced the L1 Ticket Deflector, a LangGraph-based IT service desk agent that routes tickets through a three-way state machine: auto-resolve for low-sensitivity requests, human review for sensitive ones, and a general queue for low-confidence matches. The agent's safety boundary is enforced by human-authored sensitivity labels on knowledge base articles rather than by the LLM, and on 45 hand-labeled tickets it achieved 100% decision and routing accuracy with a 55.6% sample deflection rate. The developer notes the small, clean test set makes the accuracy a sanity check rather than a benchmark, and the accompanying ROI model uses a conservative 30% deflection assumption.", "body_md": "The same tickets arrive at every IT service desk: password resets, VPN that won't connect, \"how do I install 7-Zip?\", a printer that's offline again. A widely cited industry range puts routine requests at 50–70% of L1 volume — work a script could handle, except nobody trusts a script with the tickets that actually matter.\n\nBuilding an agent that answers everything is easy. Building one that knows when to hand a ticket to a human is not. I built the second kind and open-sourced it.\n\n**Repo:** [github.com/zedxter/l1-ticket-deflector](https://github.com/zedxter/l1-ticket-deflector) (MIT)\n\nMost \"AI helpdesk\" demos optimize for deflection rate. That's the wrong target. A high deflection rate is trivial if you let the model guess.\n\nThe constraint I optimized for instead: never let the agent act on a sensitive request. Anything touching access, privileges, security, or money goes to a human with prepared context. The agent resolves the boring 30% and escalates the rest — with the relevant KB article already attached, so the human doesn't start from zero.\n\nI called it human-in-the-loop, but the honest framing is narrower: the agent has a hard boundary, and that boundary is the part worth testing.\n\nThe production version is a LangGraph state machine. Three outcomes, no cleverness:\n\n```\nclassify ──> no match / low confidence ──> queue (general L1 queue)\n   │\n   ├── sensitivity = low        ──> auto_resolve ──> notify_user\n   └── sensitivity = high/crit  ──> human_review ──> escalate ──> notify_user\n```\n\nThe classifier returns a KB article ID and a confidence score. Two gates decide the path:\n\n`0.55`, the ticket goes to the general queue. No guessing.` service_desk_l2` or `security_ops`, never to auto-resolve.\nEach KB article carries a `sensitivity` label. That's the trick: the model doesn't decide what's dangerous. The knowledge base does, and it's a human-authored field. The LLM only picks the article.\n\n``` python\ndef route_after_classify(state):\n    if not state.get(\"article_id\") or state.get(\"confidence\", 0) < CONFIDENCE_THRESHOLD:\n        return \"queue\"\n    if state.get(\"sensitivity\") in AUTO_RESOLVE_SENSITIVITY:\n        return \"auto_resolve\"\n    return \"human_review\"\n```\n\nThat's the entire safety story in six lines. Sensitive categories (access grants, license purchases, phishing reports, ransomware) can never reach `auto_resolve`, no matter what the model says.\n\nThe repo ships two implementations:\n\n`demo/`` graph/``MOCK_MODE=1` so you can run the logic without an API key.\nI kept the offline version because a demo you can't run is just a claim. Anyone can clone the repo and reproduce the numbers in ten seconds.\n\nOn 45 labeled tickets, the offline classifier scored:\n\n| Metric | Value | \n|---|---|\n| Auto-resolved | 25 | \n| Escalated to human | 18 | \n| No match → queue | 2 | \n| Decision accuracy | 100% | \n| Routing accuracy | 100% | \n| Sample deflection | 55.6% | \n\n100% accuracy is a red flag, not a trophy. It means the test set is too small and too clean. 45 tickets, hand-labeled, written by the same person who wrote the KB — of course it scores well. I'm publishing it because it's honest about what it is: a sanity check, not a benchmark. The real test is a client's messy backlog, and that number doesn't exist yet.\n\nThe deflection rate is the one I'd defend. 55.6% on this sample, but the ROI model uses a conservative **30%**, because market benchmarks for first-year deflection sit at 20–40% and I'd rather be wrong in the client's favor.\n\nConservative assumptions for a ~500-person company in DACH:\n\n| Metric | Value | \n|---|---|\n| L1 tickets / month | ~1,200 | \n| Deflection | 30% | \n| Avg handling time | 19 min | \n| IT rate (fully loaded) | €45 / hour | \n| Time saved | ~114 h/month (~0.7 FTE) | \n| Savings | ~€5,130 / month | \n| Retainer | €4,000 / month | \n| ROI | 1.3× / month (15× / year) | \n\nThe retainer is deliberately below the savings. A model where the vendor captures all the value is a model nobody signs. 1.3×/month is not a spectacular number — it's a defensible one, and defensible is what gets a pilot approved.\n\nI want to be precise about this, because \"proof of work\" usually means \"screenshot of a demo.\"\n\n**Real:** the classifier, the graph, the routing logic, the KB structure, the ROI math. Clone it, run it, read every line.\n\n**Stubbed:** the tool-calls. `itsm_create_ticket`, `directory_reset_password`, and `mdm_push_install` return strings, not real ITSM actions. Wiring them to Jira Service Management or Zammad over MCP is a pilot task, not a demo task.\n\n**Not built:** the Slack/Teams approval UI, telemetry, multilingual KB. They're on the roadmap and I'm not pretending otherwise.\n\nThree things I'd carry into the next version:\n\n`sensitivity` field on each KB article is auditable. \"The model is instructed to be careful\" is not.\nThe repo is MIT-licensed and the whole point is that you can pull it apart. If you run an IT desk and want to compare deflection on your own tickets, that's the interesting experiment.\n\n**Repo:** [github.com/zedxter/l1-ticket-deflector](https://github.com/zedxter/l1-ticket-deflector)\n\n*I build AI agents for reporting, CRM, and support workflows — with human review where it counts. Based in Potsdam, Germany.*", "url": "https://wpnews.pro/news/i-built-an-l1-ticket-deflector-that-knows-when-to-stop-langgraph-human-in-the", "canonical_source": "https://dev.to/madebyexpert/i-built-an-l1-ticket-deflector-that-knows-when-to-stop-langgraph-human-in-the-loop-4fd3", "published_at": "2026-10-03 22:57:52+00:00", "updated_at": "2026-10-03 23:07:45.060443+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-tools", "developer-tools"], "entities": ["LangGraph", "L1 Ticket Deflector", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/i-built-an-l1-ticket-deflector-that-knows-when-to-stop-langgraph-human-in-the", "markdown": "https://wpnews.pro/news/i-built-an-l1-ticket-deflector-that-knows-when-to-stop-langgraph-human-in-the.md", "text": "https://wpnews.pro/news/i-built-an-l1-ticket-deflector-that-knows-when-to-stop-langgraph-human-in-the.txt", "jsonld": "https://wpnews.pro/news/i-built-an-l1-ticket-deflector-that-knows-when-to-stop-langgraph-human-in-the.jsonld"}}