What happens when enterprise requirements hit Strands, LangGraph, and CrewAI - 45 runs measured A developer benchmarked Strands, LangGraph, and CrewAI across 45 runs to test how each agent framework handles enterprise requirements like human approval gates, audit trails, and structured output. LangGraph suspended and resumed approval gates in 6/6 runs with 0.0s resume time, Strands respected ordering but double-fired a publish call once and produced empty output three times, while CrewAI's re-run-on-feedback design caused an infinite loop of 131 LLM calls when the same rejection was returned repeatedly. The developer found that audit evidence location varies by framework control-flow model, with Strands offering the strongest trace-based audit story but also the only silent failures. What happens when enterprise requirements - human approval gates, audit trails, structured output - hit three agent frameworks? The first article measured how Strands, LangGraph, and CrewAI differ on a plain task. This one measures what happens when the task grows up: 45 more runs, same recorder proxy, same model, same tools. The headline finding: the frameworks fail differently. Strands - the model-driven one - finished with an empty output three times while exiting https://github.com/sunnydachs/agent-framework-showdown https://github.com/sunnydachs/agent-framework-showdown Read the previous article here. here https://dev.to/sunnydachs/i-built-the-same-agent-in-strands-langgraph-and-crewai-and-recorded-every-llm-call-to-see-how-4ae8 . Two walls keep coming up in developer communities when agents touch regulated work: The EU AI Act makes automatic logging and retention a legal obligation for high-risk AI systems. So past the "working demo", these are the two validations that matter. I ran them. The task: write a news digest, ask a human to approve, and only publish if approved. Publishing is a simulated destructive action that must never fire before approval. The three frameworks implement the gate differently: interrupt suspends the whole graph; Command resume=... continues it after the decision Task human input=True - the crew pauses for console feedback after the task completes Results: LangGraph suspended 6/6 runs and resumed in 0.0s - the checkpointer restores state without re-executing. Reject routes away from the publish node via an edge condition, so the gate is enforced by the graph's structure, not by the model's good behavior. Strands respected the order too: ask - check - publish 3/3, and never published on reject. But one run called publish article twice with the identical draft . The prompt was followed; the model just double-fired the CrewAI is the simplest shape: 1 call for approve, 2 for reject the feedback re-runs the task . One operational gotcha measured along the way: returning the same rejection on every prompt spins an infinite loop - 131 LLM calls with the prompt growing from 240 to 6,561 tokens. CrewAI re-runs and re-prompts on every non-empty feedback, so the repetition policy is the caller's responsibility. The regulated question is "why did the agent decide this". Because every run goes through the same recorder proxy, the traces have one shape - so I scored whether an auditor can recover the seven audit-relevant facts decision rationale, tool call order, tool arguments, model identity, and more from each framework's traces: | Framework | Rationale | Tool order | Args | Silent failures | |---|---|---|---|---| | Strands | 100% | 100% | 100% | 2 | | LangGraph | 100% | 0% | 0% | 0 | | CrewAI | 100% | 50% | 50% | 0 | Strands is model-driven, so everything the model saw and reasoned about stays in the trace - the strongest audit story of the three. The flip side is exactly those 2 silent failures. LangGraph's 0% is not a defect: its tool calls live in code, not on the wire. Read the code and you know the order; read only the trace and you don't. That is the real audit-design trade-off: where the evidence lives changes with the framework's control-flow model. Output the digest as strict JSON with exactly 4 keys summary , word count , topics , publish ready . All three frameworks hit 100% compliance, and word count matched the actual summary length in every run - putting the count inside the schema makes the model's self-verification effective. Strands ran a validate loop averaging 2 calls 4 revisions in one run . Across all 45 runs, only Strands finished with an empty output three times. The model built the complete result, handed it to the validation tool, and then emitted nothing as the final answer. The run exits 0 - it looks successful. You only catch it by reading the trace. LangGraph and CrewAI: zero. In a pipeline design, the output node IS the deliverable, so an empty answer is structurally hard to produce. For enterprise use, this is the scariest class of failure: not an error that stops the run, but a success-shaped empty result that breaks everything downstream. The empty outputs were recoverable - the model passes its full result to the tool as an argument, so the trace holds it. One model, 3 runs per cell - directional, not a definitive ranking. The human is scripted; no real UI or notification flow. The destructive action is simulated - though whether the gate held is read directly from the recorded traffic, which is the part that's solid. Everything is open. The repo README has the commands for all five experiments 72 recorded runs total, one proxy in front of every framework - that's the whole foundation : This is a personal OSS project - no warranty. Use at your own risk, and issues are welcome. Cover image: generated with a local flux-schnell pipeline.