The same tickets arrive at every IT service desk: password resets, VPN that won't connect, "how do I install 7-Zip?", a printer that's offline again. A widely cited industry range puts routine requests at 50β70% of L1 volume β work a script could handle, except nobody trusts a script with the tickets that actually matter.
Building an agent that answers everything is easy. Building one that knows when to hand a ticket to a human is not. I built the second kind and open-sourced it.
Repo: github.com/zedxter/l1-ticket-deflector (MIT)
Most "AI helpdesk" demos optimize for deflection rate. That's the wrong target. A high deflection rate is trivial if you let the model guess.
The constraint I optimized for instead: never let the agent act on a sensitive request. Anything touching access, privileges, security, or money goes to a human with prepared context. The agent resolves the boring 30% and escalates the rest β with the relevant KB article already attached, so the human doesn't start from zero.
I called it human-in-the-loop, but the honest framing is narrower: the agent has a hard boundary, and that boundary is the part worth testing.
The production version is a LangGraph state machine. Three outcomes, no cleverness:
classify ββ> no match / low confidence ββ> queue (general L1 queue)
β
βββ sensitivity = low ββ> auto_resolve ββ> notify_user
βββ sensitivity = high/crit ββ> human_review ββ> escalate ββ> notify_user
The classifier returns a KB article ID and a confidence score. Two gates decide the path:
0.55, the ticket goes to the general queue. No guessing. service_desk_l2 or security_ops, never to auto-resolve.
Each KB article carries a sensitivity label. That's the trick: the model doesn't decide what's dangerous. The knowledge base does, and it's a human-authored field. The LLM only picks the article.
def route_after_classify(state):
if not state.get("article_id") or state.get("confidence", 0) < CONFIDENCE_THRESHOLD:
return "queue"
if state.get("sensitivity") in AUTO_RESOLVE_SENSITIVITY:
return "auto_resolve"
return "human_review"
That's the entire safety story in six lines. Sensitive categories (access grants, license purchases, phishing reports, ransomware) can never reach auto_resolve, no matter what the model says.
The repo ships two implementations:
demo/`` graph/``MOCK_MODE=1 so you can run the logic without an API key.
I kept the offline version because a demo you can't run is just a claim. Anyone can clone the repo and reproduce the numbers in ten seconds.
On 45 labeled tickets, the offline classifier scored:
| Metric | Value |
|---|---|
| Auto-resolved | 25 |
| Escalated to human | 18 |
| No match β queue | 2 |
| Decision accuracy | 100% |
| Routing accuracy | 100% |
| Sample deflection | 55.6% |
100% accuracy is a red flag, not a trophy. It means the test set is too small and too clean. 45 tickets, hand-labeled, written by the same person who wrote the KB β of course it scores well. I'm publishing it because it's honest about what it is: a sanity check, not a benchmark. The real test is a client's messy backlog, and that number doesn't exist yet.
The deflection rate is the one I'd defend. 55.6% on this sample, but the ROI model uses a conservative 30%, because market benchmarks for first-year deflection sit at 20β40% and I'd rather be wrong in the client's favor.
Conservative assumptions for a ~500-person company in DACH:
| Metric | Value |
|---|---|
| L1 tickets / month | ~1,200 |
| Deflection | 30% |
| Avg handling time | 19 min |
| IT rate (fully loaded) | β¬45 / hour |
| Time saved | ~114 h/month (~0.7 FTE) |
| Savings | ~β¬5,130 / month |
| Retainer | β¬4,000 / month |
| ROI | 1.3Γ / month (15Γ / year) |
The retainer is deliberately below the savings. A model where the vendor captures all the value is a model nobody signs. 1.3Γ/month is not a spectacular number β it's a defensible one, and defensible is what gets a pilot approved.
I want to be precise about this, because "proof of work" usually means "screenshot of a demo."
Real: the classifier, the graph, the routing logic, the KB structure, the ROI math. Clone it, run it, read every line.
Stubbed: the tool-calls. itsm_create_ticket, directory_reset_password, and mdm_push_install return strings, not real ITSM actions. Wiring them to Jira Service Management or Zammad over MCP is a pilot task, not a demo task.
Not built: the Slack/Teams approval UI, telemetry, multilingual KB. They're on the roadmap and I'm not pretending otherwise.
Three things I'd carry into the next version:
sensitivity field on each KB article is auditable. "The model is instructed to be careful" is not.
The repo is MIT-licensed and the whole point is that you can pull it apart. If you run an IT desk and want to compare deflection on your own tickets, that's the interesting experiment.
Repo: github.com/zedxter/l1-ticket-deflector
I build AI agents for reporting, CRM, and support workflows β with human review where it counts. Based in Potsdam, Germany.