{"slug": "four-things-langgraph-checkpoints-won-t-do-for-multi-day-agents", "title": "Four Things LangGraph Checkpoints Won't Do for Multi-Day Agents", "summary": "A developer running multi-day agent workloads on LangGraph documented four gaps that LangGraph checkpoints do not cover, including stale session handoffs and steps that report success while leaving the target unchanged. In 17 real scenarios from the Flowness agent harness, a naive terminal success label was wrong 10 times, while an effect contract plus read-back from the target matched the true state in 9 of 9 selected operations. The write-up recommends per-task ownership records, explicit handoff sheets, and separating Attempt, Effect, Adoption and Acceptance instead of a single green tick.", "body_md": "*A longer version of this article is on [towow.ai](https://towow.ai/articles/langgraph-alternative-long-running-agents?r=devto-c04).*\n\nIf your agents lose the thread on day three, a different runtime usually will not fix it. LangChain's own comparison page puts LangGraph, Temporal and Inngest in the same group, runtimes, whose job is durable execution, streaming, human-in-the-loop and persistence ([docs](https://docs.langchain.com/oss/python/concepts/products)). A LangGraph checkpointer saves a snapshot of graph state at each super-step, per thread ([docs](https://docs.langchain.com/oss/python/langgraph/checkpointers)). What goes into that state, and whether it is true, is up to you.\n\nBelow are four gaps we hit running agent work over days. Each has the symptom, a fix, and what our own records show; the records come from Flowness, the agent harness we use in our own delivery work. The code blocks are sketches with made-up task and file names, not a library API or Flowness code. The replay sketch follows the pseudo-code in our cursor write-up, and the deploy snippet is LangGraph's own example.\n\n**Symptom.** A new session starts and either continues from a stale picture or duplicates work the old session still holds.\n\n**Fix.** Keep a handoff record in graph state or a store ([persistence docs](https://docs.langchain.com/oss/python/langgraph/persistence)), and call a check like this at the start of whichever node runs first when a new session picks up the thread. A sketch, not a library API:\n\n```\nhandoff = {\n    \"in_flight\": [{\"task\": \"t-17\", \"owner\": \"session-a\", \"step\": \"tests\",\n                   \"artifact\": \"branch/perm-tests\", \"blocked_on\": None}],\n    \"next_waits_on\": \"t-17 review\",\n    \"refs\": [\"task_packet@v3\", \"branch/perm-tests\", \"report/latest\"],\n}\n\ndef resume_gate(handoff):\n    for ref in handoff[\"refs\"]:\n        if not exists_and_current(ref):   # packet version, branch, readable report\n            raise NeedsRepair(ref)        # fix the sheet; don't continue from memory\n    for t in handoff[\"in_flight\"]:\n        if activity_belongs_to(t[\"task\"], t[\"owner\"]):\n            skip_or_attach(t)             # owner still working: don't start a twin\n```\n\nCheck ownership per task, not per project line. \"Someone is on the permissions work\" does not tell you whether this test task is taken.\n\n**Our record.** Our handoff sheet lists in-flight tasks, owners, artifacts, blockers and what the next step waits on, and the receiver checks every reference first ([write-up, in Chinese](https://towow.net/articles/harness-long-task-handover)). In a 2 September 2026 review of one rebuild, 29 of 32 tasks were recorded as successful while one requirement had no task carrying it all the way ([write-up, in Chinese](https://towow.net/articles/flowness-goal-across-handoffs)). The sheet keeps the work; the goal needs its own object, inherited by every stage.\n\n**Symptom.** A step prints success, exits 0, and the target is unchanged.\n\n**Fix.** A checkpoint records that a node returned. It cannot see the repository, service or file the node claims to have changed. For each step with an outside effect, declare what to observe and where, then read it back from the target, not from the process that made the change:\n\n```\neffect = Effect(target=\"repo:main\", expect=\"perms.py contains check_access\")\n\nresult = run_step()                       # Attempt\nobserved = target_readback(effect)        # fresh read from the target: Effect\nrecord(attempt=result, effect=observed)   # Adoption and Acceptance get their own labels\n```\n\nReport Attempt, Effect, Adoption and Acceptance separately instead of one green tick.\n\n**Our record.** In 17 real Flowness scenarios a naive terminal label was wrong 10 times. In 9 selected operations, stdout and exit code each matched the true state in 4; effect contract plus read-back matched in 9 ([study](https://towow.net/en/research/when-done-did-not-happen)). These are selected sets, not a general failure rate.\n\n**Symptom.** After a crash, a restart double-counts or double-acts.\n\n**Fix.** Assume every step runs twice. LangGraph's docs say replay re-executes nodes after the chosen checkpoint, so LLM calls and API requests fire again ([time travel](https://docs.langchain.com/oss/python/langgraph/use-time-travel)). Interrupts re-run their node, so side effects before an interrupt should be idempotent ([interrupts](https://docs.langchain.com/oss/python/langgraph/interrupts)). In `exit` durability mode, intermediate state is not saved, so a mid-run crash is not recoverable ([checkpointers](https://docs.langchain.com/oss/python/langgraph/checkpointers)).\n\nFor your own side effects, save the result and the dedup evidence in one atomic write, and advance the progress marker last:\n\n```\npending, end = read_after_cursor()\nwith lock():\n    state = load()\n    for offset, key in pending:\n        if offset > state.high_water:     # already-counted offsets are skipped\n            state.counts[key] = state.counts.get(key, 0) + 1\n    state.high_water = max([state.high_water] + [o for o, _ in pending])\n    atomic_replace(state)                 # counts and high-water mark together\nadvance_cursor(end)                       # last\n```\n\n**Our record.** With synthetic offsets 40, 80 and 120, a crash before the cursor moves leaves counts of 2; the restart rereads three signals and ends at 2 + 0 + 0 + 1 = 3, where naive re-adding gives 2 + 3 = 5. The write-up covers a fixed, append-only local file, a lock among cooperating writers, and recovery from process exit, not power loss ([write-up, in Chinese](https://towow.net/articles/harness-cursor-recovery)).\n\n**Symptom.** You pause the system and work keeps arriving. Or you ship a fix and in-flight runs change behavior.\n\n**Fix.** Write down what a pause stops: new dispatch, in-flight work, review and fix lanes, and coordinator patrols are set separately. On resume, read what arrived during the pause before restarting any dispatch; resuming everything at once can restart work someone already holds.\n\nFor deploys, LangGraph's backward-compatibility guide says the latest graph is applied to every thread, including threads resuming from a checkpoint, whereas some workflow engines pin a run to its starting code version. Renaming or removing a node while threads are paused at it breaks the resume. The guide recommends stamping a behavioral version on state at thread start and branching on it ([docs](https://docs.langchain.com/oss/python/langgraph/backward-compatibility)):\n\n``` python\ndef intake(state):\n    return {\"flow_version\": state.get(\"flow_version\", 2)}  # new threads are stamped 2\n\ndef after_triage(state):\n    return \"policy_check\" if state.get(\"flow_version\", 1) >= 2 else \"respond\"  # old threads default to 1\n```\n\n**Our record.** In a 15 September pause, two in-flight executors finished their current section, new dispatch went to zero, and review and fix lanes stayed on ([write-up, in Chinese](https://towow.net/articles/harness-long-task-handover)).\n\nLangChain's page lists Temporal and Inngest beside LangGraph as runtimes, and the Deep Agents SDK and Claude Agent SDK as harnesses. That grouping is LangChain's; this post does not test them. The four fixes above are about how you design state, handoffs and checks, so they carry to any runtime.\n\n*Drafted with AI assistance, based on our project records, fact-checked before publishing.*", "url": "https://wpnews.pro/news/four-things-langgraph-checkpoints-won-t-do-for-multi-day-agents", "canonical_source": "https://dev.to/natureblueee/four-things-langgraph-checkpoints-wont-do-for-multi-day-agents-1gom", "published_at": "2026-10-10 19:01:15+00:00", "updated_at": "2026-10-10 19:16:27.184602+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-tools", "mlops", "developer-tools"], "entities": ["LangGraph", "LangChain", "Temporal", "Inngest", "Flowness"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/four-things-langgraph-checkpoints-won-t-do-for-multi-day-agents", "markdown": "https://wpnews.pro/news/four-things-langgraph-checkpoints-won-t-do-for-multi-day-agents.md", "text": "https://wpnews.pro/news/four-things-langgraph-checkpoints-won-t-do-for-multi-day-agents.txt", "jsonld": "https://wpnews.pro/news/four-things-langgraph-checkpoints-won-t-do-for-multi-day-agents.jsonld"}}