{"slug": "the-llm-didn-t-win-everywhere-and-that-s-what-made-the-project-interesting", "title": "The LLM Didn't Win Everywhere — and That's What Made the Project Interesting", "summary": "A developer closed a key stage of ops-triage-ai, an operational triage system that combines a deterministic baseline, a local LLM, and a hybrid policy with human review. Evaluated on a frozen held-out set of 70 synthetic tickets in a single run, the deterministic baseline still beat the local LLM on overall risk accuracy, 95.7% versus 91.4%, and the regression was published alongside the gains. The project reports a hybrid latency of roughly 6 to 7 seconds and 13 published limitations, and has not been validated under real production traffic.", "body_md": "I closed an important stage of `ops-triage-ai`, an operational triage\n\nsystem that combines a deterministic baseline, a local LLM, and a hybrid\n\npolicy with human review.\n\nThe most interesting result was not simply \"the LLM was better.\"\n\nEvaluation on a frozen held-out set of 70 synthetic tickets, in a single\n\nrun. Deterministic baseline versus the local LLM:\n\nAnd the point that matters: the deterministic baseline still won on\n\noverall risk accuracy — 95.7% against the LLM's 91.4%.\n\nThe LLM did not win everywhere. The regression is published alongside the\n\ngains, with nothing smoothed over.\n\nInstead of:\n\n\"How do we replace rules with AI?\"\n\nthe architecture answers:\n\n\"How do we combine different behaviors in a way that is safe, observable,\n\nand auditable?\"\n\nIn the final benchmark:\n\nReview rate describes operations here — how many cases asked for review —\n\nnot review precision/recall. The held-out set has no independent ground\n\ntruth for that.\n\nSynthetic dataset, in English, with 70 examples. A single official run,\n\nwith no measurement of LLM variance. The system has not been validated\n\nunder real production traffic. Hybrid latency sits around 6 to 7 seconds:\n\nfine for asynchronous triage, not for a critical synchronous path. The\n\nfull details — official metrics, trade-offs, and the 13 published\n\nlimitations — are in the [project case](https://marcelotaparelli.com.br/en/projects/ops-triage-ai/).\n\nA useful AI system does not need to trust the model blindly.\n\nSometimes the best architecture comes precisely from understanding where\n\nthe model is better — and where it is not.", "url": "https://wpnews.pro/news/the-llm-didn-t-win-everywhere-and-that-s-what-made-the-project-interesting", "canonical_source": "https://dev.to/marcelotaparelli/the-llm-didnt-win-everywhere-and-thats-what-made-the-project-interesting-a1a", "published_at": "2026-09-15 09:33:21+00:00", "updated_at": "2026-09-15 09:39:17.522918+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-agents", "ai-tools", "mlops"], "entities": ["ops-triage-ai"], "alternates": {"html": "https://wpnews.pro/news/the-llm-didn-t-win-everywhere-and-that-s-what-made-the-project-interesting", "markdown": "https://wpnews.pro/news/the-llm-didn-t-win-everywhere-and-that-s-what-made-the-project-interesting.md", "text": "https://wpnews.pro/news/the-llm-didn-t-win-everywhere-and-that-s-what-made-the-project-interesting.txt", "jsonld": "https://wpnews.pro/news/the-llm-didn-t-win-everywhere-and-that-s-what-made-the-project-interesting.jsonld"}}