From memory to behavior: skills, user profiles, and the consolidation loop Neo4j Labs' meta-knowledge-graph repository adds a consolidation layer that converts accumulated agent session learnings into reusable skills and user profiles, addressing the problem where twelve sessions on the same hook-pipeline debugging produce twelve separate learnings that similarity search returns only five of. The layer routes checked learnings through two paths: task learnings become a Skill with SkillVersion records and pitfalls via DERIVED_FROM and INFORMED_BY, while user preferences become a UserProfile with profile versions tracked through CONSOLIDATED and FOLDED_LEARNING. Before contributing, each learning passes a safety screen for disguised instructions and credential material plus a consistency judge, with ambiguous conflicts surfaced through the project_gate_audit MCP tool and human overrides via project_resolve_learning. From memory to behavior: skills, user profiles, and the consolidation loop Graph ML and GenAI Research, Neo4j 16 min read Turn accumulated learnings into procedures and user profiles, then feed them back into the agent Part 3 of the self-learning multi-agent series. Part 1 covered the tools and hooks . Part 2 introduced the semantic layer and memory - Coauthored with Firat Tekiner https://medium.com/u/d95593087c94 Adding memory to an agent is fairly straightforward. Store something useful from a session, retrieve it when a similar task comes up, and include it in the next prompt. The agent no longer has to discover everything from scratch. After a few hundred sessions, though, the agent has a different problem. Imagine that twelve sessions produce twelve learnings about debugging the same hook pipeline. One explains how session deduplication works. Another records a cooldown that prevented a background job from running. A third captures the tool error that started the investigation. A similarity search might return five of them, leaving the agent to work out how they fit together each time. The individual learnings can all be useful while the collection remains difficult to use. We need a way to turn those fragments into a procedure. And once we have that procedure, we should stop injecting the fragments alongside it. This is the problem that the consolidation layer in the meta-knowledge-graph repository https://github.com/neo4j-labs/meta-knowledge-graph addresses. We want the debugging sessions to produce a reusable skill, with the steps that worked and the failures worth avoiding. Along the way, the agent may also learn that the user prefers concise explanations or a particular workflow. Those observations belong in a personal profile that follows the user across projects. The diagram shows two paths from the same session history. Task learnings become a procedure the agent can reuse, while user preferences become a profile that follows the person across projects. Both paths start with checked learnings. Let’s follow the task path first, then the user profile path. The upper half represents task knowledge. Each : Learning links back to its origin sessions through FROM SESSION and joins a reusable task category through TAGGED WITH . A : TaskPattern brings related learnings together across sessions, while OBSERVED IN connects that category to the sessions where the work occurred. A : Skill links to the learnings it was built from through DERIVED FROM , and to error learnings that shaped its pitfalls through INFORMED BY . Its : SkillVersion records preserve proposals and their outcomes as the procedure evolves. The lower half represents personal preferences. Each : User has a : UserProfile , which links to the user-scoped learnings it has incorporated through CONSOLIDATED . Profile versions preserve the summary over time, with FOLDED LEARNING recording which facts were folded into each version. Both paths keep the reusable guidance connected to its sources, so a later correction can be traced to the skill or profile it affects. Checking and grouping learnings Before a learning contributes to a skill or profile, it passes two checks. A safety screen looks for instructions disguised as facts, attempts to weaken safeguards, and credential material. A consistency judge then compares it with nearby learnings in the same project and scope, merging restatements, rejecting unsupported changes, or superseding older knowledge. Ambiguous conflicts remain visible through the project gate audit MCP tool, and humans can override gate decisions through the project resolve learning MCP tool. Those checks evaluate individual learnings. To assemble a procedure, we also need to know which learnings belong to the same task. “The injection hook deduplicates on session id” and “run the Stop hook twice to check the cooldown branch” describe different parts of debugging the hook pipeline. Similarity between their texts may not be enough to bring them together. The extractor therefore assigns a task pattern while it still has the session context. This is a short, reusable caption of at most six words, such as hook pipeline debugging, cypher schema migration, or flaky test triage. It describes the kind of work, giving learnings from different sessions a common grouping key. At extraction time, the task pattern caption is stored as a property on the learning. Resolving it to a : TaskPattern node happens later, inside the background skill consolidation service, before the service asks the LLM to propose a skill. The hooks invoke this service on Stop and, where supported, SessionEnd . These events give it an opportunity to run; they do not force consolidation after every conversation. By default, new distillation work proceeds only when the project has at least four pending eligible learnings and the twenty-four-hour cooldown has passed. Eligible learnings are approved, project-scoped, carry a task label, and are not already sources of an active skill or an outstanding proposal. Once those checks pass, the service loads eligible learnings together with the source learnings of existing live skills. Previously assigned pattern links are reused. For each unassigned learning, it normalizes the label by lowercasing it, replacing runs of punctuation or whitespace with a single space, and trimming the result. Thus, Hook-Pipeline Debugging and hook pipeline debugging resolve to the same pattern through an exact match within the project. For labels without an exact match, the service uses embeddings of the short pattern labels/captions. It compares the incoming label’s vector with the project’s existing pattern vectors using cosine similarity and reuses the nearest pattern if its score reaches MKG TASK PATTERN SIMILARITY THRESHOLD, which defaults to 0.80. Otherwise, it creates a new : TaskPattern node and stores the label and its embedding there. New patterns immediately join the matching pool, so paraphrases within the same batch can also converge on one node. If embeddings are unavailable, matching falls back to normalized text only, which can leave differently worded labels in separate groups. The implementation batches embedding requests for unassigned labels before matching them; exact matches still take precedence over vector similarity. The service then writes learning - :TAGGED WITH - pattern and links the pattern to the learning’s origin sessions through pattern - :OBSERVED IN - session . Our deduplication note, cooldown observation, and verification step can now meet under hook pipeline debugging, with access to the session evidence behind them. Group membership follows these shared pattern links. The fuzzy matching therefore happens while preparing the groups, before a skill is proposed, and the links persist even if a group produces no proposal in that run. For example, through my use of the plugin, the system identified these task patterns. Turning task learnings into a reusable skill The skill service uses those grouped learnings to assemble a procedure with four required sections: when to use it, the steps to follow, pitfalls, and how to verify the result. Instead of leaving the next agent to piece together five retrieved fragments, it gives the agent a procedure grounded in the earlier work. We also collect errors and their known resolutions as learnings, so the resulting skill can explicitly describe pitfalls and fixes. Session links let the service find relevant error learnings from earlier attempts and include them when assembling the procedure. Some systems use a PostToolUse hook to inject relevant learnings immediately after a tool call. Here in our system, errors and resolutions enter the memory pipeline and contribute to later recall and skill consolidation. The following diagram shows the implementation behind both consolidation paths. The skill and profile services run in the background after sessions, with thresholds and cooldowns controlling when they do work. By default, the skill service requires at least four pending eligible learnings and waits twenty-four hours between runs . It is triggered by conversation hooks Stop and, where supported, SessionEnd , not a cron job. If no conversation produces another hook event, the service does not run, even after the cooldown expires. Learnings sharing a task pattern form a group. A new skill needs at least two learnings in its group; if the group includes a source of an existing live skill , one new learning can justify a patch. Consolidation gives experience time to accumulate. A single observation may be useful on its own; several related learnings can explain how to carry out a task. The service periodically considers that growing collection as conversations finish. It runs through conversation hooks, so an idle project stays idle: the passage of time alone does not start a consolidation run. Return to our hook debugging example. - the deduplication note explains why context might appear twice, - the cooldown observation explains why a job might not run, - and the verification step explains how to check the behavior. Together, they can become a debugging procedure that connects symptoms, checks, and expected results. The model decides whether the group supports a useful procedure, adds something to an existing one, or should remain as individual learnings. That procedure can improve as the agent does more work. A later session might uncover a missing prerequisite or a better way to verify the result. The new learning can update the existing skill, preserving what still works and correcting what has changed. The aim is to build a small collection of procedures that become more useful with experience, rather than a new document after every session. Turning a memory into a skill also changes its influence: a description of past work becomes guidance for future actions. Before making it available, the service checks that the proposed steps are grounded in the source learnings and screens the procedure for unsafe instructions. Screened skills become available automatically by default, with an optional human review step for teams that want to approve procedures before agents use them. The original learnings remain available while a proposal is being prepared or reviewed. Once the skill becomes available, it takes over as the reusable guidance, and its source learnings remain connected to it in the graph. Those links make later corrections possible: if a source learning turns out to be wrong or outdated, the system can identify the affected skill and flag it for revision. A skill can also be retired, allowing its still-valid source learnings to return to ordinary recall. Discovering skills incrementally Plugin skills are typically advertised to the agent in the system prompt through a catalog of names or slugs and short descriptions. Their full instructions can be loaded later, but the agent still receives the catalog up front. As the collection grows, even that catalog takes up context before the agent knows which procedures it will need. Graph-backed skills add another level of deferral: the agent can discover relevant entries through MCP, then expand only the ones it needs. The two tools divide that work: - skill search task finds candidate procedures for the current task and returns slugs with one-line descriptions, without their bodies. - skill fetch slug expands a selected result into its full procedure, including the source learning references. For the next hook debugging task, the agent can search for “debug hook pipeline,” inspect the returned candidates, and fetch a matching procedure. If the results do not fit, it can refine the search before loading any instructions. Context grows in steps: first the discovery tools, then a small set of candidates, then the selected procedure. Discovery can also start automatically when the user submits a prompt. The prompt hook reuses the text embedding already computed for learning recall and compares it with the embeddings of the project’s approved skills. It injects up to five relevant skill slugs whose cosine similarity is at least 0.8, controlled by MKG SKILL INJECT MIN SIMILARITY. The agent can then expand a suggested slug with skill fetch. If no skill clears the threshold, or the prompt cannot be embedded, this path adds no skill suggestions. It never injects procedure bodies. The session-start hook can still provide a compact catalog of live skill slugs. The two shortcuts are independently configurable: MKG SKILL CATALOG INJECT=0 disables the startup catalog, while MKG SKILL PROMPT INJECT=0 disables prompt-time suggestions. Both are enabled by default. Explicit skill search remains available whether or not either hook suggests a skill, so the agent can discover procedures without their slugs in its initial context. Because the procedures live in the graph, any agent with access to these MCP tools can use them without a harness-specific skill file. Once a skill is live, its source learnings are marked consolidated and excluded from ordinary context recall. This keeps the twelve debugging fragments from being injected again alongside the procedure they produced. New, unconsolidated observations can still enter context through learning recall and later improve the skill. The original records and embeddings remain searchable by the extractor and consistency gate for deduplication, comparison, and inspection. The task path now ends with a procedure the agent can discover when similar work comes up. The other consolidation path uses what those sessions taught us about the user. Consolidating the user’s personal profile While the sessions were teaching the agent how to debug the pipeline, they may also have taught it how the user likes to work. A preference for concise explanations belongs across projects, whereas a cooldown in this repository does not. Approved user-scoped learnings therefore go through a separate consolidation service that maintains a compact personal profile. Previously, approved user facts were folded into the shared :SystemPrompt. With multiple users, that mixes personal preferences into shared operating instructions. Consolidation now maintains a separate :UserProfile attached to each user, leaving the shared base prompt unchanged. The diagram shows two users whose facts are folded independently. User A already has a third profile version, with the two earlier versions archived, while user B has just received a first profile. The base prompt is the same in both sessions; only the appended section differs. The service combines the existing profile with newly approved user facts. For example, repeated sessions may establish that the user prefers concise explanations and wants a proposed approach before implementation. Consolidation folds those preferences into a small set of adaptations the agent can apply across projects. By default, this runs when more than five approved user facts are pending and the twenty-four-hour cooldown has passed. The model edits the existing profile rather than appending every fact. The result is capped at twelve bullets and approximately fifteen hundred characters; oversized replies are discarded, and the outgoing version is archived. Even after the gate checks them, the facts are presented as untrusted data to summarize. Corrections also need to reach the profile. If a contributing fact is retracted, the profile is marked stale. A stale profile bypasses the normal threshold and cooldown so the service can remove the retracted guidance promptly. Once a preference has been folded into the profile, its source learning is marked consolidated and excluded from ordinary recall, just like the sources of a live skill. Its embedding and record remain available for deduplication and inspection. An edited user fact can return to the profile backlog so the summary stays aligned with the underlying memory. The return path is automatic. At session start, inject system prompt.py loads the shared base prompt and appends that user’s profile under User adaptations. The agent begins with the person’s preferences already available, while task procedures remain available to fetch when relevant. The next debugging session can therefore start with both outcomes from the diagram: a reusable procedure assembled from previous work and a personal profile that shapes how the agent works with this user. Further sessions can improve the procedure and update the profile without requiring the agent to reconstruct either from a growing collection of fragments. Summary Memory alone leaves the agent with a growing collection of fragments. Consolidation turns those fragments into two kinds of guidance. Task learnings that share a task pattern become a skill with - the steps that worked, - the pitfalls to avoid, - and a way to verify the result. User-scoped facts become a personal profile that follows the person across projects, with one profile per user and a shared base prompt that stays unchanged. Neither result is loaded into context wholesale. Skills are discovered incrementally: a prompt-time match or an explicit skill search call returns slugs with short descriptions, and the agent expands only the procedure it needs with skill fetch. The profile is compact by construction, capped at twelve bullets, and appended at session start. Once a learning has been folded into a skill or a profile, it leaves ordinary recall, so the agent receives the procedure instead of the fragments that produced it. Everything stays in the graph, with its provenance. A skill links to the learnings it was derived from and to the error learnings that shaped its pitfalls, and those learnings link to the sessions where the work happened. A profile links to the facts it consolidated, and each archived version records which facts it folded. When a source learning is retracted or superseded, the system can find the skill or profile it affected and flag it for revision or repair, instead of leaving stale guidance in place. None of this requires a person to run it. Hooks give the services an opportunity to run after each conversation, thresholds and cooldowns decide whether enough new material has accumulated, the gate checks every candidate before it is reused, and a safety screen checks every proposed skill. Screened skills go live on their own by default, with an optional human review step for teams that want it. The agent begins each session with what earlier sessions taught it, and the graph records how it came to know it. From memory to behavior: skills, user profiles, and the consolidation loop https://medium.com/neo4j/from-memory-to-behavior-skills-user-profiles-and-the-consolidation-loop-9406fd1aa9dd was originally published in Neo4j Developer Blog https://medium.com/neo4j on Medium, where people are continuing the conversation by highlighting and responding to this story.