You spend three evenings with an AI on a billing pipeline. On the fourth, you open a fresh chat, paste a summary the model wrote the night before, and ask it to continue. Within five messages it proposes Prisma. You rejected Prisma on day one, for a reason that took twenty minutes to establish.
The summary was accurate. It was also missing the one thing the new session needed: what is already settled, and why.
Ask a model to "summarize our conversation" and it will compress. Compression favors the narrative: what we discussed, what we built, roughly where we ended. It drops the parts that look like noise, and some of that noise is load-bearing.
Three kinds of information tend to disappear:
A new session reads that summary as the full truth. Anything absent is treated as undecided, so it gets decided again, often differently.
I wrote about the basic habit in Before You Close That ChatGPT Tab, Run This One Command First. This article is about the structure that makes the habit reliable, because the wording of the request decides whether you get a summary or a working document.
The reason to hand off at all is that long conversations degrade. In Lost in the Middle: How Language Models Use Long Contexts, the authors found that models perform best when relevant information sits at the beginning or end of the input, and worse when it sits in the middle. In a chat with eighty turns, your early constraints live in that middle.
Anthropic's engineering team describes the same pressure from the agent side. In Effective context engineering for AI agents, they note that as the number of tokens in the window grows, the model's ability to recall information from it decreases (they call this context rot), and describe compaction as summarizing a conversation nearing the limit so it can be reinitialized. Their example of what a good compaction keeps is telling: architectural decisions, unresolved bugs, and implementation details, while discarding redundant tool output.
Notice what that list is not. It is not a recap of the conversation. It is state and commitments. That distinction is the whole design of a good handoff.
Author's Comment: When I read a handoff, I do not check whether it sounds complete. I check whether a stranger could continue the work without asking me a single question. If the answer is no, something binding is missing.
The Session Handoff Document prompt forces the model to fill six fixed sections. Each one blocks a specific failure.
auth.ts", not "made progress on authentication". Specificity here prevents the new session from redoing work.
Sections 4 and 5 do the real work. They are two sides of the same line. One lists what is closed. The other lists what is open. A session that knows both has a boundary to work within.
In the prompt's sample output, section 4 reads like this:
- Runtime & ORM: Node.js 22 + TypeScript (strict) + Drizzle ORM.
Rejected Prisma due to cold-start latency in serverless workers.
- No synchronous downstream calls: the webhook endpoint must ACK
within 150ms; heavy lookups happen in the async worker.
- Locking: use SELECT ... FOR UPDATE SKIP LOCKED instead of
adding Redis/BullMQ, to avoid a new infrastructure dependency.
Each line carries a decision and a reason. The reason matters as much as the decision. A bare "no Prisma" invites the model to ask why and reopen the case. "Rejected Prisma due to cold-start latency" gives it the logic to respect the choice, and to notice if a new requirement genuinely changes it.
I call these decision locks. They are the cheapest insurance in the whole document, and the first thing a casual "summarize this" request leaves out.
A handoff document alone is only half the transfer. The prompt ends by producing a ready-to-paste bootstrap block that tells the next session how to read it:
Read the Session Handoff Document above. Do not re-debate the
Confirmed Decisions. First, confirm your understanding in 3 concise
bullet points, then begin executing Priority #1 under Next Steps.
The three-bullet confirmation is a cheap alignment check. If the model restates the state wrong, you catch it before it writes a line of code. If it restates it correctly, you have a short, verifiable record that the handoff landed.
The prompt takes three inputs: project_domain, next_session_goal, and handoff_depth. They change the shape of the output more than you would expect.
A coding handoff should capture file paths, schema choices, and active bug traces. A writing handoff should capture voice rules, the outline, and the thesis guardrails. A strategy handoff needs trade-offs and risks. The handoff_depth presets cover these directly, from a comprehensive state injection to a compact snapshot tuned for minimal tokens.
Setting next_session_goal is the most underrated of the three. It tells the model what the next session will do first, so it can prioritize which details to preserve in full and which to compress.
Practical Pitfall Avoidance Guide: Read the handoff once before you copy it. If the thread contained a wrong assumption that was never corrected, the model will serialize that assumption as fact. Fix the premise first, then generate the handoff. A clean handoff of a derailed chat is a polished mistake.
Skip it for short, stateless tasks. A single lookup, a one-off rewrite, or a quick translation does not carry state worth transferring, and a handoff adds overhead with no return.
It earns its place on work that spans sessions: refactors, long drafts, research with a growing list of excluded sources, and planning where early decisions constrain later ones.
A handoff only helps if you can find it when the next session starts. Keep them in a prompt manager such as Prompt Vault, one entry per project, and paste the latest one at the top of each new chat.
Before your next long session ends, run the handoff prompt with your domain and next goal filled in. Read sections 4 and 5 first. If a decision you remember making is not there, add it before you close the tab.