A longer version of this article is on towow.ai.
If your agents lose the thread on day three, a different runtime usually will not fix it. LangChain's own comparison page puts LangGraph, Temporal and Inngest in the same group, runtimes, whose job is durable execution, streaming, human-in-the-loop and persistence (docs). A LangGraph checkpointer saves a snapshot of graph state at each super-step, per thread (docs). What goes into that state, and whether it is true, is up to you.
Below are four gaps we hit running agent work over days. Each has the symptom, a fix, and what our own records show; the records come from Flowness, the agent harness we use in our own delivery work. The code blocks are sketches with made-up task and file names, not a library API or Flowness code. The replay sketch follows the pseudo-code in our cursor write-up, and the deploy snippet is LangGraph's own example.
Symptom. A new session starts and either continues from a stale picture or duplicates work the old session still holds.
Fix. Keep a handoff record in graph state or a store (persistence docs), and call a check like this at the start of whichever node runs first when a new session picks up the thread. A sketch, not a library API:
handoff = {
"in_flight": [{"task": "t-17", "owner": "session-a", "step": "tests",
"artifact": "branch/perm-tests", "blocked_on": None}],
"next_waits_on": "t-17 review",
"refs": ["task_packet@v3", "branch/perm-tests", "report/latest"],
}
def resume_gate(handoff):
for ref in handoff["refs"]:
if not exists_and_current(ref): # packet version, branch, readable report
raise NeedsRepair(ref) # fix the sheet; don't continue from memory
for t in handoff["in_flight"]:
if activity_belongs_to(t["task"], t["owner"]):
skip_or_attach(t) # owner still working: don't start a twin
Check ownership per task, not per project line. "Someone is on the permissions work" does not tell you whether this test task is taken.
Our record. Our handoff sheet lists in-flight tasks, owners, artifacts, blockers and what the next step waits on, and the receiver checks every reference first (write-up, in Chinese). In a 2 September 2026 review of one rebuild, 29 of 32 tasks were recorded as successful while one requirement had no task carrying it all the way (write-up, in Chinese). The sheet keeps the work; the goal needs its own object, inherited by every stage.
Symptom. A step prints success, exits 0, and the target is unchanged.
Fix. A checkpoint records that a node returned. It cannot see the repository, service or file the node claims to have changed. For each step with an outside effect, declare what to observe and where, then read it back from the target, not from the process that made the change:
effect = Effect(target="repo:main", expect="perms.py contains check_access")
result = run_step() # Attempt
observed = target_readback(effect) # fresh read from the target: Effect
record(attempt=result, effect=observed) # Adoption and Acceptance get their own labels
Report Attempt, Effect, Adoption and Acceptance separately instead of one green tick.
Our record. In 17 real Flowness scenarios a naive terminal label was wrong 10 times. In 9 selected operations, stdout and exit code each matched the true state in 4; effect contract plus read-back matched in 9 (study). These are selected sets, not a general failure rate.
Symptom. After a crash, a restart double-counts or double-acts.
Fix. Assume every step runs twice. LangGraph's docs say replay re-executes nodes after the chosen checkpoint, so LLM calls and API requests fire again (time travel). Interrupts re-run their node, so side effects before an interrupt should be idempotent (interrupts). In exit durability mode, intermediate state is not saved, so a mid-run crash is not recoverable (checkpointers).
For your own side effects, save the result and the dedup evidence in one atomic write, and advance the progress marker last:
pending, end = read_after_cursor()
with lock():
state = load()
for offset, key in pending:
if offset > state.high_water: # already-counted offsets are skipped
state.counts[key] = state.counts.get(key, 0) + 1
state.high_water = max([state.high_water] + [o for o, _ in pending])
atomic_replace(state) # counts and high-water mark together
advance_cursor(end) # last
Our record. With synthetic offsets 40, 80 and 120, a crash before the cursor moves leaves counts of 2; the restart rereads three signals and ends at 2 + 0 + 0 + 1 = 3, where naive re-adding gives 2 + 3 = 5. The write-up covers a fixed, append-only local file, a lock among cooperating writers, and recovery from process exit, not power loss (write-up, in Chinese).
Symptom. You the system and work keeps arriving. Or you ship a fix and in-flight runs change behavior.
Fix. Write down what a stops: new dispatch, in-flight work, review and fix lanes, and coordinator patrols are set separately. On resume, read what arrived during the before restarting any dispatch; resuming everything at once can restart work someone already holds.
For deploys, LangGraph's backward-compatibility guide says the latest graph is applied to every thread, including threads resuming from a checkpoint, whereas some workflow engines pin a run to its starting code version. Renaming or removing a node while threads are d at it breaks the resume. The guide recommends stamping a behavioral version on state at thread start and branching on it (docs):
def intake(state):
return {"flow_version": state.get("flow_version", 2)} # new threads are stamped 2
def after_triage(state):
return "policy_check" if state.get("flow_version", 1) >= 2 else "respond" # old threads default to 1
Our record. In a 15 September , two in-flight executors finished their current section, new dispatch went to zero, and review and fix lanes stayed on (write-up, in Chinese).
LangChain's page lists Temporal and Inngest beside LangGraph as runtimes, and the Deep Agents SDK and Claude Agent SDK as harnesses. That grouping is LangChain's; this post does not test them. The four fixes above are about how you design state, handoffs and checks, so they carry to any runtime.
Drafted with AI assistance, based on our project records, fact-checked before publishing.