Why Our CLI Deadlock Detector for Multi-Agent LLMs Timed Out: A Post-Mortem on Naive Jaccard Heuristics A developer built a stateless CLI tool, the Asynchronous Semantic Deadlock Checker, to detect semantic deadlock among collaborating LLM agents by computing pairwise Jaccard similarity over agent output tokens and flagging divergence below a 0.15 threshold. During QA stress tests with 20 or more active agents, the tool collapsed with an unhandled timeout, which the post-mortem attributes to the naive whitespace tokenization and O(N^2) pairwise similarity matrix. When designing orchestrations for autonomous multi-LLM workflows, one of the most frustrating failure modes is semantic deadlock —where multiple agents fall into circular reasoning, context divergence, or redundant conversational loops. In an attempt to catch these divergence states early in automated continuous-integration workflows and long-running daemons, we built a standalone, stateless CLI tool: the Asynchronous Semantic Deadlock Checker . The goal was simple: consume execution traces, calculate semantic alignment across all collaborating agents, and trigger recovery interventions within a strict 10-second SLA. However, during high-concurrency QA stress tests involving 20+ active agents, the tool collapsed with an unhandled timeout. Here is the complete technical breakdown of the architecture, the production code, the profiling post-mortem, and the architectural anti-patterns exposed by this failure. To ensure zero daemon overhead and easy integration into CI/CD pipelines, pre-commit hooks, and cron-based monitoring scripts, we designed the checker around a strictly stateless Unix-philosophy pipeline: stdin , containing recent execution traces, agent identifiers, operational states, and raw text outputs. looping , stuck , repeating , contradicting . stdout , detailing deadlock verdicts, offending agent IDs, and automated mitigation directives FORCE RESET CONTEXT AND CONVERGE . graph TD subgraph InputStage "Input Processing" A "Agent Logs: JSON Stream via STDIN" -- B "Token Extraction: Naive Whitespace Splitting" end subgraph MetricStage "Pairwise Analysis" B -- "O N^2 Matrix Generation" -- C "Jaccard Distance Matrix" C -- "Threshold < 0.15" -- D "Divergence Flag" C -- "Outlier Distance" -- E "Culprit Identification" end subgraph DecisionStage "Resolution & Recovery" D -- F "Deadlock Decision Engine" E -- F F -- "Emit Payload" -- G "Recovery Instructions / Reset Trigger" end Below is the complete Python implementation deployed into the test environment. python import sys import json from collections import defaultdict def tokenize text : if not text: return set return set str text .lower .split def jaccard similarity set a, set b : if not set a or not set b: return 0.0 intersection = len set a.intersection set b union = len set a.union set b return intersection / union if union 0 else 0.0 def main : try: input data = sys.stdin.read if not input data.strip : print json.dumps {"error": "Empty input"}, ensure ascii=False return logs = json.loads input data except Exception as e: print json.dumps {"error": f"Invalid JSON: {str e }"}, ensure ascii=False return agents = defaultdict list for entry in logs: aid = entry.get "agent id", "unknown" agents aid .append entry agent latest = {} agent tokens = {} for aid, history in agents.items : latest = history -1 agent latest aid = latest agent tokens aid = tokenize latest.get "output", "" agent ids = list agent latest.keys n = len agent ids similarity matrix = {} deadlock suspects = divergence detected = False Pairwise similarity calculation: O N^2 for i in range n : id a = agent ids i tokens a = agent tokens id a for j in range i + 1, n : id b = agent ids j sim = jaccard similarity tokens a, agent tokens id b similarity matrix f"{id a}-{id b}" = sim if sim < 0.15: divergence detected = True for aid, latest in agent latest.items : state = latest.get "state", "idle" if state in "looping", "stuck", "repeating", "contradicting" : deadlock suspects.append aid Culprit isolation via mean similarity thresholding if not deadlock suspects and divergence detected and n 1: avg sims = {} for id a in agent ids: sims = jaccard similarity agent tokens id a , agent tokens id b for id b in agent ids if id a = id b avg sims id a = sum sims / len sims if sims else 1.0 if avg sims: culprit = min avg sims, key=avg sims.get deadlock suspects.append culprit suspect set = list set deadlock suspects result = { "status": "DEADLOCK DETECTED" if divergence detected or deadlock suspects else "NORMAL", "divergence detected": divergence detected, "culprit agents": suspect set, "similarity matrix": similarity matrix, "correction instruction": { "target agents": suspect set, "action": "FORCE RESET CONTEXT AND CONVERGE", "message": "Semantic deadlock or context divergence detected. Refocus on the primary objective and abandon circular arguments." } } print json.dumps result, ensure ascii=False, indent=2 if name == " main ": main 💡 For immediate deployment: The complete source code suite ZIP for this architecture is available on Gumroad https://phenox.gumroad.com/l/ozvtlz for $0+ Pay What You Want . In our QA stress-testing environment, we simulated an orchestration topology with 24 concurrent agents , where each agent generated long-form diagnostic dumps, chain-of-thought traces, and structured code snippets totaling several thousand tokens per turn. Under these conditions, execution terminated with the following unhandled exception: TimeoutExpired: Command ' 'python3', 'deadlock checker.py' ' timed out after 10 seconds. The script contained no infinite loops while True or unbounded recursion . The timeout was triggered entirely by an algorithmic and memory management bottleneck. For $N$ agents, calculating the full similarity matrix requires evaluating $\frac{N N - 1 }{2}$ unique pairs. When $N = 24$, this yields 276 pairwise evaluations. Furthermore, during the culprit isolation phase: sims = jaccard similarity agent tokens id a , agent tokens id b for id b in agent ids if id a = id b The script recalculates Jaccard similarities from scratch rather than referencing the previously populated similarity matrix hash table, adding an extra $N N - 1 = 552$ evaluations. While 828 set operations might appear trivial on paper, the underlying payloads turned this into a massive bottleneck. Calling str text .lower .split across 24 distinct multi-thousand-token strings creates hundreds of thousands of intermediate string instances. Converting these arrays into hashable Python set structures triggers significant allocation overhead. During nested set intersections set a.intersection set b , the Python runtime repeatedly hashes and traverses large lookup tables in memory. The sheer volume of ephemeral objects triggered continuous Garbage Collection GC pauses, destroying CPU throughput inside the single-threaded CPython interpreter. Beyond raw compute bottlenecks, the core premise proved semantically defective. The Jaccard index measures purely lexical, surface-level token overlap: $$J A, B = \frac{|A \cap B|}{|A \cup B|}$$ In complex LLM dialogue systems, agents frequently enter circular arguments while employing completely different phrasing, terminology, or syntactical structures e.g., "We must redesign the REST contract" vs. "The API endpoint interface requires refactoring" . Because the lexical intersection was near zero, the algorithm reported false-positive context divergence sim < 0.15 . In an attempt to patch these false positives, we had layered additional outlier-detection heuristics onto the script, expanding execution complexity and accelerating the performance collapse. This failed experiment highlights critical architectural constraints when building runtime monitors for multi-agent systems: text-embedding-3-small , mini-LM models, or quantized on-device encoders paired with Cosine Similarity or Approximate Nearest Neighbor ANN search. Treating complex semantic arbitration as a one-shot CLI text-manipulation script was an oversimplification of the problem space. The Asynchronous Semantic Deadlock Checker failed its operational requirements as a sub-10-second CLI tool and was officially shelved. However, this failure provided an unambiguous architectural boundary: Reliable multi-agent oversight cannot be achieved through ad-hoc string comparisons. Monitoring multi-agent consensus requires an embedded vector-state pipeline, streaming sliding-window telemetry, and dedicated vector-distance indexing. When designing guards against LLM deadlocks, do not settle for lexical shortcuts. Build your monitoring infrastructure on genuine embedding spaces from day one. If this engineering log saved your production server and your sanity , consider supporting our architecture on GitHub Sponsors.