{"slug": "codesmithi-a-textbook-anatomy-of-agents-five-waves-of-evolution", "title": "CodeSmithi: A Textbook Anatomy of Agents: Five Waves of Evolution", "summary": "A developer published a technical anatomy of the CodeSmith agent framework (v0.5.0), using its reference loop implementation in DefaultAgentExecutor::run_inner to illustrate the ReAct reason-act-observe cycle and the argument that an agent's action trajectory cannot be collapsed into a single longer static prompt. The writeup frames the field around five waves of agent evolution — loop, autonomy, multi-agent, and harness engineering — and notes that environmental feedback from each executed step becomes new context that later reasoning depends on.", "body_md": "Source version of [CodeSmith](https://github.com/camilesing/CodeSmith): `v0.5.0` (commit `3a74c82f`). All paths are relative to the repo root; line numbers refer to this version.\n\nIntended audience: readers new to Agent engineering who want a coordinate system of \"loop–autonomy–multi-agent–evolution.\"\n\nThe prologue told of an invisible war, then tossed out a word: harness. Over the next twenty-odd installments we will plunge headlong into CodeSmith's organs — caches, handles, compaction, approval gates. But before the scalpel comes out, this installment wants to pause and dissect the word \"Agent\" itself: what the loop looks like, how many levels autonomy comes in, whether multi-agent can be trusted, and why in 2026 every engineering team is talking about harness engineering. These are the common knowledge of this craft, owned by no single project; and CodeSmith's source will keep surfacing to confirm them — including the passage of code pasted below, the closest thing to a textbook in the entire repository.\n\nThe core abstraction of the modern Agent is the ReAct loop: the model first reasons about the current state and the actions available, executes one action (search, query a database, run code), the environment returns an observation, and the model reasons its next round from that observation. It sounds bland, but it carries a less-than-obvious corollary: **an Agent's action trajectory cannot be reduced to one longer static answer**.\n\nThe reason is that action changes what information is available. In pure-reasoning mode, if the context holds no information about \"whether this code compiles,\" the model cannot \"think it up\" — it can only guess. In the loop, by contrast, the model runs the compile command, the compiler's error output becomes a new fact in the context, and subsequent reasoning is built on those new facts. An Agent that has run twenty steps owes its final judgment to the environmental feedback of every one of those steps — feedback that, before the actions took place, existed nowhere at all. So \"replacing the loop with a longer prompt\" holds only when the task involves no interaction with the environment.\n\nInside CodeSmith's engine sits a loop implementation that reads like a textbook. Not the 16,000-line `host_executor` on the production path, but the reference implementation in the framework crate, `DefaultAgentExecutor::run_inner` (`crates/agent/src/executor/mod.rs:140`), whose module-header comment bills itself as \"The LangChain `AgentExecutor` analog\". Strip away the peripheral details and its skeleton looks like this:\n\n``` js\n// crates/agent/src/executor/mod.rs:140 (excerpt)\nlet mut step: u32 = 0;\nloop {\n    if step >= max_steps {\n        callback.on_complete(&StopReason::MaxSteps).await;\n        return Ok(StopReason::MaxSteps);\n    }\n    // ...assemble the MessageRequest (model, message history, system, tool catalog)...\n    let stream = client.create_message_stream(request).await?;\n    let (content, _stop_reason) = accumulate_stream(stream).await?;\n\n    // Persist the assistant turn.\n    history.push(Message { role: \"assistant\".to_string(), content: content.clone() });\n\n    // Collect tool calls (preserve order).\n    let tool_uses: Vec<(String, String, serde_json::Value)> = content\n        .into_iter()\n        .filter_map(|block| match block {\n            ContentBlock::ToolUse { id, name, input, .. } => Some((id, name, input)),\n            _ => None,\n        })\n        .collect();\n\n    if tool_uses.is_empty() {\n        return Ok(StopReason::NoToolCalls);\n    }\n\n    // Execute each tool sequentially and feed the result back as a\n    // `role:\"user\"` `ToolResult` block (Anthropic/OpenAI-compat shape).\n    for (id, name, input) in tool_uses {\n        let result = match tools.get(&name) {\n            Some(tool) => tool.run(input.clone()).await,\n            None => Err(ToolError::NotAvailable { /* no tool named '{name}' */ .. }),\n        };\n        history.push(Message {\n            role: \"user\".to_string(),\n            content: vec![ContentBlock::ToolResult { tool_use_id: id, /* ... */ }],\n        });\n    }\n    step += 1;\n}\n```\n\nThis code is doing exactly one thing — ask the model, collect tool calls from the reply, execute them and stuff the results back into the history under the user role, until the model stops asking for tools. \"Thinking\" is the assistant message, \"acting\" is the ToolUse block, \"observing\" is the backfilled ToolResult — ReAct's three beats, mapped word for word onto the message structure. Two details deserve a second look: tool results return to the history under the identity of `role:\"user\"` (this foreshadows Article 10's status-bar design), and calling a nonexistent tool is not a crash but a `NotAvailable` error result — fed back to the model, so that it corrects itself.\n\nThe stop conditions are gathered into a single enum (`crates/agent/src/callback/mod.rs:23`):\n\n```\npub enum StopReason {\n    /// The model produced an assistant turn with no tool calls — the run is\n    /// finished.\n    NoToolCalls,\n    /// The step budget (`max_steps`) was exhausted mid-tool-loop.\n    MaxSteps,\n    /// The run aborted with an error.\n    Error(String),\n    /// The run was cancelled (user/external interruption). Distinct from\n    /// [`StopReason::Error`] so the host can surface \"cancelled\" rather than\n    /// \"error\" — mirrors production's `TurnOutcomeStatus::Interrupted`.\n    Interrupted,\n}\n```\n\nFour values tell the whole story of the loop's exits: the model finished speaking, the steps ran out (`max_steps` defaults to 50), an error occurred, the user interrupted. \"User interruption\" gets its own variant rather than being folded into Error so that the interface can truthfully say \"cancelled\" instead of \"error\" — honest stop conditions and honest error reports are one and the same virtue.\n\nThe step-loop skeleton of the production `HostAgentExecutor` (`crates/agent-runtime/src/engine/host_executor.rs:2642`) matches the reference implementation, but roughly ten guardrails hang on before and after each step: cancellation checkpoints, system-prompt snapshot refreshes, compaction, capacity preflight, LSP diagnostics flush, the loop guard... Those guardrails are the protagonists of this series' second half; for now, they stay offstage.\n\nIn industry usage, the word \"Agent\" suffers obvious marketing inflation. Lay the common misuses of the concept out on the table, and the boundaries of every discussion that follows become much clearer:\n\n| Claim | What it actually is | One-sentence puncture | \n|---|---|---|\n| \"Autonomous agent completed the search\" | A single API call | A system that keeps no state, chooses no actions based on observations, and has no stop condition is not an Agent | \n| \"Autonomous workflow\" | A fixed pipeline dressed up | All the decisions are hard-coded; the model only fills in the blanks | \n| \"The Agent gets better as it works\" | Long-chain error propagation | One wrong file edit contaminates every subsequent operation; autonomy has nothing to do with it | \n| \"Experiment succeeded, tests passed\" | Environmental hallucination | The \"it's done\" narrative had no real execution behind it — more dangerous than text hallucination | \n| \"The foundation model has this capability\" | Capability misattribution | The gains from good retrieval, dedicated interfaces, and heavy retries get booked to the model's weights | \n| \"The Agent decided to…\" | The autonomy myth | Choosing an action among given tools does not mean it \"wants\" or \"decides\" | \n\nOne touchstone for telling the real from the fake: when the environment returns an unexpected result, can the system change its sequence of actions? If it can, it has earned the right to talk about loops; if it cannot, it is nothing more than a fill-in-the-blank exam with a temperature setting. The touchstone cuts both ways — against the things we build ourselves too. CodeSmith's loop guard (Article 10) exists precisely because the distance between a system that \"can change its action sequence\" and one \"destined to repeat the same action\" is a single guardrail.\n\nSomeone has run a systematic ablation study on an Agent's context: keep the full baseline, then remove one component at a time as the controls — tool definitions, tool execution results, the reasoning process, message history (the system prompt is the identity definition; ablate it and even running the test becomes meaningless, so it does not take part). The conclusions deserve to be memorized line by line:\n\nThe experiment's core insight compresses into one sentence: **the context determines what the Agent can see, and the Agent can only make decisions from what it sees.** And the components are not equivalent — the measure is whether the information a component carries can be rebuilt from elsewhere. One more finding matters even more to engineering practice: the typical failure under a mutilated context is not an error exit but a flawless-looking answer — \"produced a reply\" is not \"completed the task.\" All the context engineering of Articles 12 through 15, and the checklist by which the improvement plan audits itself against the textbook, stand on this foundation.\n\nSWE-agent made a far-reaching discovery: the same foundation model performs worlds apart on codebase-modification tasks under a plain shell interface versus under a purpose-designed Agent-Computer Interface (ACI). How the interface presents file contents, how it formats edit commands, how it returns error messages — each of these bears on whether the model can operate code effectively. Put differently — **an Agent's capability is not an attribute of the model weights alone, but a joint product of the model and the environment's interface.**\n\nThis is the academic rendering of the prologue's \"the difference is the harness.\" Take the harness apart and three groups of components fall out:\n\nFit these three groups onto CodeSmith's 21 crates and you will find them lined up in tidy ranks: `agent-runtime`'s Constitution and prompts and `tool-impls`' fifty weapons belong to the cognitive interface; `agent`/` providers`' message history and sandbox belong to the execution environment; `execpolicy`'s command review, side-git snapshots, and loop guard belong to audit and constraint. Every installment this series has ahead of it is about one block of these three groups.\n\n\"Autonomy\" is not a have-it-or-not property but a continuous variable, divisible into at least five levels. Level 1: the developer specifies each action step by step, and the model only fills in text. Level 2: the model chooses actions from a given tool set — the base mode of most Agent systems today. Level 3: the model can revise the plan, abandoning the original path when the environment returns a surprise. Level 4: the model can propose its own subgoals and decompose them. Level 5: the model can examine the task's goal and its evaluation criteria themselves — \"is this task even worth doing?\" The first four levels answer \"how to get the task done\"; the fifth pushes the question to \"whether the task itself stands.\"\n\nBy this scale, CodeSmith lives between levels 2 and 3: the model freely chooses tools and arguments (level 2), and the loop guard and the capacity controller can force a VerifyAndReplan that resets the trajectory when it drifts (a passive level 3). Level 5 is, for now, a luxury for any production system.\n\nThe loop itself has evolved a family of variants. Along three dimensions — what is saved, what is read, and what triggers the next round — at least five distinct forms can be told apart:\n\n| Form | What it saves | Echo in this series | \n|---|---|---|\n| ReAct (reactive) | No extra memory maintained across steps | The reference implementation above | \n| Reflexion (reflective retry) | After failure, generates self-reflection, stores it in external memory, reads it back together next time | Article 23's experience distillation | \n| LATS (search tree) | Each step is a node in a search tree; multiple candidate paths backtrack by value function | — | \n| Voyager (skill accumulation) | Successful action sequences are encoded into reusable skill programs and filed into a skill library | Article 2's skills; Article 23 | \n| MemGPT (layered memory) | A layered memory system with active read/write, managing history beyond the window | Article 8's virtual memory | \n\nThe differences among the five forms are, in essence, differences of memory strategy: ReAct is \"no memory,\" Reflexion is \"remember only on failure,\" Voyager is \"remember only on success,\" MemGPT is \"memory itself needs paging.\" In Article 8 you will see the bloodline connecting CodeSmith's VarHandle to MemGPT — that paper's metaphor of choice was the operating system's virtual memory.\n\nWhen a task passes from one Agent to another, the delegation has to be contracted explicitly, or it will fail at some boundary. A complete delegation contract has at least eight clauses: objective, tool permissions, forbidden actions, resource ceilings, abort conditions, output format, lines of responsibility, and a renegotiation mechanism for when the environment changes. Leave out the resource ceiling and the delegate may burn compute without end; leave out the abort condition and a subtask that has already drifted from its goal will simply keep producing useless results.\n\nTo anyone too optimistic about multi-agent implementations, I would like to throw three buckets of cold water:\n\nSo multi-Agent value demands specific conditions: the task must genuinely decompose, the subtasks must be independent enough, and the coordination mechanism must be able to handle dependencies. Article 6 shows how CodeSmith builds work crews under these constraints — every member with its own prompt, its own model, its own git worktree — which is, in essence, the responsibilities and resource boundaries from the \"eight elements of the delegation contract\" turned into configuration items.\n\nPull the camera back from any single system, and AI application engineering has traced a clean arc over the past few years:\n\n```\nGraph engineering    — Agent loops, deterministic programs, and human approval organized into an explicit execution graph\n└── Loop engineering — sustained autonomous operation across turns: who spots the next thing, when to verify, when it counts as done\n    └── Harness engineering — context and tool interfaces, constraints, verification, feedback loops, error recovery\n        └── Context engineering — systematically managing everything the model can see\n            └── Prompt engineering — optimizing the natural-language instructions fed to the model\n```\n\nNotice that these five waves do not replace one another — they nest: prompt engineering is a subset of context engineering, context engineering a subset of harness engineering — and the individual Agent loop is precisely one node in the execution graph. Each layer widens the engineer's field of attention beyond the one before it.\n\nWhy does this arc point beyond the model? LangChain's practice on Terminal Bench 2.0 (a benchmark that evaluates Agents completing complex tasks in a terminal environment) supplies a forceful footnote: the score climbed from 52.8% to 66.5%, vaulting from outside the top thirty on the leaderboard into the top five — **what was swapped was not the model, it was the harness**. The concrete measures included letting the Agent automatically check its own execution results, detecting whether it had fallen into a repetitive loop, and refining its thinking strategy. As the capabilities of the various models converge and cease to be the decisive differentiator, competitive advantage migrates to the engineering practices outside the model. That sentence is the reason this entire series exists.\n\nOn engineering principles, Anthropic gathers the lessons of successful Agents into three: keep it simple (direct API calls beat complex frameworks; every additional layer of abstraction is a fresh blind spot for future debugging); keep it transparent (planning, logs, and decision trajectories stay visible — an error inside a black box can be neither located nor corrected); and design the tool interface well — design it from the Agent's point of view, and where misuse comes easily, make the error impossible by design. Manufacturing has a term of art for the third: poka-yoke (mistake-proofing), out of the Toyota Production System — the notched corner of a SIM card makes inserting it backwards impossible, and a microwave with its door not properly shut will not heat. These three will keep echoing through every installment to come: the Constitution's layering is simplicity, the status bar is transparency, `execpolicy`'s three-valued Decision is poka-yoke.\n\nGather this installment's contents into six sentences:\n\nThe anatomy chart is drawn. Next comes the first cut, beginning with the most valuable item on this checklist of distrust: the 100× price tag of a single byte.", "url": "https://wpnews.pro/news/codesmithi-a-textbook-anatomy-of-agents-five-waves-of-evolution", "canonical_source": "https://dev.to/dogeking/codesmithi-a-textbook-anatomy-of-agents-five-waves-of-evolution-5ea6", "published_at": "2026-10-02 16:01:00+00:00", "updated_at": "2026-10-02 16:08:21.214734+00:00", "lang": "en", "topics": ["ai-agents", "artificial-intelligence", "large-language-models", "ai-tools", "developer-tools"], "entities": ["CodeSmith", "DefaultAgentExecutor", "LangChain", "AgentExecutor", "Anthropic", "OpenAI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/codesmithi-a-textbook-anatomy-of-agents-five-waves-of-evolution", "markdown": "https://wpnews.pro/news/codesmithi-a-textbook-anatomy-of-agents-five-waves-of-evolution.md", "text": "https://wpnews.pro/news/codesmithi-a-textbook-anatomy-of-agents-five-waves-of-evolution.txt", "jsonld": "https://wpnews.pro/news/codesmithi-a-textbook-anatomy-of-agents-five-waves-of-evolution.jsonld"}}