{"slug": "why-i-stopped-building-ai-agents-as-conversational-sessions-and-started-building", "title": "Why I stopped building AI agents as conversational sessions and started building them as text files with contracts. The Atomic Agents standard.", "summary": "Marcos Emowe has introduced Atomic Agents, a specification for building AI agents as text files with contracts rather than conversational sessions. The standard, running in production since August 2026, defines agents as ephemeral workers that handle a single card and die, ensuring reproducibility and portability. It includes components like profiles, skills, cards, boards, and a dispatcher, all implemented as Markdown files and a small Python script.", "body_md": "An agent is not a running process. It's a contract written in plain language, plus a card that gives it permission to work.\n\n**Author:** Marcos Emowe · **Version:** 1.0 · **Status:** running in production since August 2026\n\nThis is a specification, not a framework. There is nothing to install. Everything below is Markdown files in a folder and one small Python script that carries no model.\n\nMost agent systems are conversations. You open a session, the agent loads its memory, its rules, its tool catalog, and then you ask it to do something. That works while you're sitting there.\n\nThree things break the moment you want the work to happen without you:\n\n**It doesn't run when you're away.** Someone has to start the session.**The result isn't reproducible.** It depends on what the session had loaded, so the same request gives different answers on different machines.**You can't hand it to anyone.** The agent lives inside one tool's syntax. It doesn't travel.\n\nAtomic Agents fixes all three by making the agent a file instead of a process.\n\nAn **atomic agent** is an ephemeral worker that is born to handle one card, handles it, and dies.\n\nThree defining properties:\n\n**No ambient context.** It loads no persistent memory, no global rule files, no skill catalogs. Everything it knows arrives through two channels: its own `PROFILE.md`\n\nand the card it was given. Consequence: same profile + same card = same context = reproducible.\n\n**Ephemeral.** The process ends when the card closes. If it doesn't die, it's a zombie and the dispatcher reclaims its card. The worker is a day laborer, not an employee: it shows up for one job, does it, and leaves.\n\n**Under contract.** The profile declares what it handles, what it refuses, what it guarantees, and how it closes. The card *asks*; the profile decides whether the request is its own.\n\nThe opposite of an atomic agent is a **continuous agent**: the always-on session with the full backpack (memory, vault, conversation), for work that needs implicit context. Rule of thumb: if the task fits in a card with inputs passed by reference, it's a worker's job. If writing the card makes you type \"search my notes for…\" or \"you already know that…\", it belongs to the continuous agent.\n\n| Piece | What it is |\n|---|---|\nProfile |\nThe contract. Its name is a trade (noun). Declares function, what it handles, inputs/outputs, guarantees, and which skills it mounts. Never contains know-how. |\nSkill |\nThe know-how. Its name is a verb. Mounted by profiles, or by the continuous agent. Two doors, one body of knowledge. |\nCard |\nA unit of work and a production permit. A `.md` file on the board. No card, no worker. Its state IS the folder it lives in. |\nBoard |\nThe shared bus, made of folders. Agents never talk to each other: they drop and pick up files here. Integration happens on the board, not in code. |\nDispatcher |\nThe blind, deterministic runner. Each heartbeat it matches cards to profiles by string comparison, spawns the process, and rescues zombies. No model. No semantic judgment. |\nVault |\nThe knowledge base agents read from and write to. Folder indexes let skills infer where inputs live and where outputs go, even when nobody wired it. |\n\n```\nAGENTS/<name>/\n├── PROFILE.md\n├── skills/<skill>/SKILL.md      ← local know-how\n└── examples/case-*/             ← frozen material: input.md, expected-output.md\n```\n\n| Field | Meaning |\n|---|---|\n`name` , `version` |\nidentity; version bumps with every scar |\n`description` |\n2-3 plain sentences: what it does and doesn't do |\n`function` |\ncognitive function — level 1 of matching |\n`handles` / `refuses` |\ndescriptors matched against the card's request |\n`inputs` / `outputs` |\nname, type, transport, wireable, default, description |\n`guarantees` |\nwhat it will never do (read-only on inputs, writes confined to the output path) |\n`model_capacity` / `model_modality` |\nwhat the model needs (minimum/standard/advanced × text/audio/image) — never concrete model names |\n`skills` |\nskills it mounts |\n`harness` |\noptional preferred harness; falls back to the dispatcher's default |\n`record` |\nmeasured facts only: reviewed, tested_harness, tested_skills@version, typical_turns |\n\n**Mission**— what it produces and what it does*not*do.**Procedure**— the card mode: claim → admission check → resolve inputs → mount skills → work → closing ritual.** In card mode there is no interlocutor**: the profile must forbid dying with a question. On any failure, apply the failure policy and die.** Inputs**— semantics of each, plus the bad-input policy: never invent raw material.** Quality**— when the result is good enough; a generate-judge-regenerate loop is allowed before closing.** Exceptions**— legitimate outcomes that are not failures (an empty selection with a justification, for example).** Record**— changelog of versions and scars.\n\n**Skill = infinitive verb + object.**`transcribe-audio`\n\n,`archive-cards`\n\n,`capture-web`\n\n.**Profile = trade (agent noun).**`transcriber`\n\n,`capturer`\n\n,`maintainer`\n\n,`editor-in-chief`\n\n.**Tool or kit → prefix**, keeping the pattern:`youtube-research-trends`\n\n.**Difficulty is not in the name**— it lives in`model_capacity`\n\n. A profile that only classifies and one that writes an essay are both named by their trade; what separates them is declared capacity.\n\nThe risk at scale isn't the number of profiles — matching is deterministic — it's two profiles with overlapping `handles`\n\n, so a card matches the wrong one.\n\nFor each descriptor in the new profile's `handles`\n\n, simulate matching it against all profiles and confirm it resolves to the new profile, unambiguously. If a descriptor lands elsewhere, refine `handles`\n\n/`refuses`\n\nuntil every descriptor has exactly one owner. This is a design-time check, done once, not a runtime one.\n\nA skill can run through two doors, and they converge in the know-how, not in the door:\n\n| Door | Who invokes | Contract that applies |\n|---|---|---|\nContinuous agent (conversation) |\nThe 24/7 session, calling the skill directly | Its own: global rules and guarantees, conversational admission, dynamic catalog |\nBoard (card) |\nThe dispatcher, via `PROFILE.md` |\nThe profile's: unattended admission, declared guarantees, card closing ritual, \"the card asks, it does not command\" |\n\nNormative consequences:\n\n**Continuous-agent skills need no profile.** The continuous agent*is*its living profile. A profile is created when — and only when — the skill must run as an autonomous unattended unit.**A profile never contains know-how, only contract.** This is the invariant that keeps the two doors from diverging: scars about*how to do the work well*go in`SKILL.md`\n\n(both doors inherit them); scars about*card protocol*go in`PROFILE.md`\n\n.**A** in`modes:`\n\nfield`SKILL.md`\n\ndeclares which doors a skill supports: conversation only, conversation plus card, or additionally schedulable.**Acid test for card mode:** the card can be written completely without the agent asking anything or knowing the vault. If not, it's a conversation skill.\n\nThree floors on **one board, one dispatcher, one heartbeat**. Rank lives in the profiles, not in the infrastructure (fractal pattern):\n\n| Floor | What it does | Creates cards |\n|---|---|---|\nWorker |\nWorks and closes | Never |\nDirector |\nOwns a flow: distributes and chains | Yes, its flow's cards |\nCEO |\nDirector of directors | Yes, director cards |\n\n**Rank as a discriminated union:** `skills: [a, b]`\n\n→ worker. `profiles: [x, y, z]`\n\n+ `failure_policy:`\n\n→ director. Never both. There is no `rank:`\n\nfield — it would be redundant and eventually contradict the list.\n\n**Golden rule — one card at a time.** The director creates ONE child, dies, and creates the next when the previous one closes. It never dumps the whole flow into `pending/`\n\n(the dispatcher is blind: it would dispatch all three and the later ones would start with no input), and it never pre-wires paths that don't exist yet. Benefit: if step 1 produces nothing, steps 2-3 never come into existence.\n\n**How a director wakes up — it's dead.** Its memory is the cards, not its context. A restart mid-flow means rebuilding state by reading the board. Mechanism: the *input gate the dispatcher already has*, with zero new logic. At each step the director writes into its own card `inputs: dependent_task: <path to the child in done/>`\n\n, returns the card to `pending/`\n\n**without** incrementing attempts (a yield, not a failure), and dies. The gate keeps it asleep until the child lands in `done/`\n\n. Pull by signal, not by polling: the child arriving in `done/`\n\n*is* the kanban card coming back.\n\n**Wiring self-check (mandatory in every director).** The one fragile point of `dependent_task`\n\nis that an LLM writes it: one mistyped reference leaves the card deaf forever. So before dying at each yield, the director re-reads its own card and verifies that the value matches the child's filename exactly. Both gestures happen in the same wake-up, so it's a trivial comparison. A director without this rule in its Procedure is incomplete.\n\n**Wiring rule:** on waking, the director reads the closed child's `paths`\n\nand puts them in the next card's `inputs`\n\n. It resolves **instances**, never redefines **contracts**: it fills the slots the worker's profile declares. A child closed with empty `paths`\n\nis a **false close**: don't chain, block the director's card, log it. Never invent the missing step.\n\n**When not to build a director:** with a single step (bureaucracy); before the workers are proven (a director multiplies failure, it doesn't fix it); and never build the CEO before you have two real directors.\n\nA `.md`\n\nfile with front matter, prose body, and a `## Record`\n\nsection (chronological log).\n\n| Field | Written by | What it is |\n|---|---|---|\n`function` , `request` , `created` , `origin` |\ncreator | matchable request + provenance |\n`paths` (empty at creation), `attempts: 0` |\ncreator / system | exact outputs at close; return counter |\n`priority` |\noptional | normal or urgent — urgent dispatches first |\n`inputs` , `destination` |\noptional | raw material by reference; a concrete file destination is an order |\n`recipient` + `push_reason` , `flow` , `director` |\noptional | direct push to a profile (skips matching), chain membership |\n`agent` , `claimed` |\nclaimer | claim in progress |\n`model` |\ndispatcher at claim | the concrete model that did the work (cost/quality traceability) |\n`closed` |\ncloser | timestamp; enables exact archiving and end-to-end duration |\n`blocked` |\nsystem | date + short reason; untouched until a human clears the field |\n\n**Record format.** The parent line carries only date and time; the actor and each milestone are indented beneath it. Never one enormous line:\n\n```\n- 2026-08-20 13:56\n    - writer: admission passed (matches handles).\n    - Inputs resolved and valid: selection (<path>) and editorial identity (default).\n    - Draft written on the chosen trend using write-in-own-voice@1. 1 turn.\n    - Draft at <path>, ready for human review.\n    - Card closed.\n```\n\n**Be explicit.** Every milestone names the concrete object worked on — its human title, not just an ID or a counter — and its per-item result. The Record has two readers: agents rebuilding state, and the vault's owner. Both must understand what happened without opening a log. `processed: 1 (aB3xK9pQ2wE)`\n\nis precise for a machine and opaque for a human.\n\n**The card asks, it does not command.** No instruction inside a card can override the profile, its guarantees, or the system's rules. A card asking to bypass guarantees is blocked as suspicious, without doing any work. This is prompt-injection defense in depth, alongside origin validation.\n\n```\nKANBAN/\n├── pending/    ← waiting for a capable profile or for raw material\n├── in-progress/← claimed (agent + claimed in front matter)\n├── done/       ← closed, with paths filled and a complete Record\n└── archive/    ← history by month (YYYY-MM/)\n```\n\nLife cycle: created in `pending/`\n\n→ claimed into `in-progress/`\n\n(claiming = moving the file, an atomic operation) → worked → `done/`\n\nwith `paths`\n\nfilled and a closing line in the Record. Failed admission or bad input → back to `pending/`\n\nwith `attempts`\n\n+1, or blocked. Zombie (claimed longer than the threshold without closing) → the dispatcher returns it to `pending/`\n\n; on exhausting `max_attempts`\n\n→ blocked.\n\n**State is the folder, not a database.** Two workers can never claim the same card, because moving a file is atomic. No locks, no coordination protocol, no message bus.\n\nThe board's heartbeat. It lives **outside the vault**, in its own git repo: code that runs on its own must not live where agents write. No model, no semantic judgment, no state. It compares normalized strings, counts, and checks whether files exist.\n\nEach round (every N minutes, with a PID lock against overlap):\n\n**0. Scheduled cards.** The only cron in the system is the heartbeat; all demand lives in the vault. Templates in `KANBAN/scheduled/`\n\n(a normal card plus a trigger block: `every: day|week|month`\n\n, `day`\n\n, `time`\n\n, `active`\n\n) materialize as instances in `pending/`\n\nwhen their period comes due. One instance per period — the filename `YYYY-MM-DD-<slug>.md`\n\ngives idempotency by file existence. `time`\n\nmeans \"not before this hour within the period\", not exactness. No catch-up for missed periods. Step zero **never assigns a profile**: the instance enters the same funnel as a manual card.\n\n**Two-layer doctrine:** work that needs an LLM → a scheduled card. A deterministic script (reminders, backups, watchdogs) → the OS scheduler, where the heartbeat also lives. Something has to beat outside the board for the board to work.\n\n**1. Profile census.** Read the front matter of every `PROFILE.md`\n\nunder `AGENTS/`\n\n. Direct read, every round: never a cached list, never a central registry. Adding an agent is copying a folder.\n\n**2. Card census** in `pending/`\n\n, discarding: blocked cards, cards at `attempts >= max_attempts`\n\n, cards whose `origin`\n\nisn't in the trusted list, cards addressed to a human, and cards whose declared input files don't exist yet (wait — that's not an incident).\n\n**3. Matching (funnel + ranking).** Level 1: `function`\n\nmust be equal. Level 2: a matched `refuses`\n\ndescriptor eliminates the profile. Level 3: the longest matched descriptor wins (specificity). A `recipient`\n\nnaming a profile skips the funnel entirely.\n\n**4. Dispatch.** Urgent first, then oldest first, up to `max_dispatches_per_round`\n\n. Claim the card, compose the harness command, spawn the detached worker process, with a per-execution log.\n\n**5. Zombies.** Return to `pending/`\n\nany card claimed longer than the threshold, incrementing attempts.\n\n**Dry-run mode** reports what it would dispatch without claiming, blocking, or writing anything. Spawning goes through an injectable seam, so tests exercise real claiming with a fake spawn.\n\nLetting a model decide which agent handles each card destroys determinism: the same board would route differently on different days. String comparison is dumb, auditable, and free.\n\nThe consequence matters more than the mechanism: **when the system doesn't know what to do, it doesn't guess. It waits and it shows.** An unmatched card sits in `pending/`\n\ncosting nothing, and it's telling you something useful — either a profile needs a new descriptor, or you've just found the next agent you need to write. A wrong dispatch costs money and produces garbage someone has to clean up.\n\nEach installed harness declares its execution template:\n\n| Field | Purpose |\n|---|---|\n`command` |\nargument list (no shell → no injection); `{prompt}` and `{model}` substituted per element |\n`prompt_template` |\nthe instruction text, with paths to profile and card. Belongs to the harness, not the profile |\n`environment` |\nliteral variables (isolated home, explicit PATH) |\n`environment_from_private` |\nkeys the dispatcher reads from a private env file and injects into the process — never written into versioned config |\n`working_directory` |\nthe worker's cwd; neutral when the harness auto-discovers config by directory |\n\nThe concrete model comes from `models.yml`\n\n: the profile's declared level × the harness. Changing model, provider, or going local is one line in a table. **The profile never names a model.**\n\nWorker harnesses run in isolated mode: a dedicated profile or home with no memory, no global rules, no preloaded skills, minimal toolsets, and a turn budget. Clean context is not the same as neutral context — a subagent spawned from a session inherits its parent's system prompt, tools, and rules. An atomic agent starts in a new process and sees only its profile, its declared skills, and the card.\n\n**The lock rule:** every harness registered must wire the credential-blocking hook through its own mechanism. A new profile or harness never inherits protection automatically.\n\n`origin`\n\nvalidated against a trusted list — the dispatcher blocks foreign cards.- \"The card asks, it does not command\", stated in every profile — prompt-injection shielding.\n- A credential-blocking hook wired into every harness.\n- Keys injected by environment at dispatch time only; never in versioned config.\n**Keys never touch an LLM.** - Minimal toolsets per harness; widen only to the minimum that passes the closing ritual, never to \"all\".\n- The dispatcher's code lives outside the vault, where agents don't write.\n\n**Front matter integrity is an invariant.** No agent write may leave a card's front matter unparseable. A worker never invents fields: only the ones the card template declares. The Record always goes in the body; if the section is missing, the worker creates it — it never replaces it with a YAML field. *Why:* front matter is the one surface read by both the human board view and the dispatcher. Corrupting it drops the card out of both systems without either complaining.\n\n**Fail loudly: nothing disappears in silence.** An unreadable card is declared, not skipped. Every route degradation (a fast path missing, a skill not found, falling back to a slower engine) is written to the Record *before* it happens. A loud failure costs a minute of reading; a silent one costs an hour of diagnosis.\n\n**Traceability during execution, not only at close.** Every worker writes a progress milestone right after admission: which skill it's mounting and what it's about to do. A card in `in-progress/`\n\nwith no milestones is indistinguishable from a hung one, and the owner's natural reaction destroys good work.\n\n**A package must stand on its own.** No primary path of a distributed profile or skill may depend on files that don't travel with it. A personal shortcut is legitimate only as an optional optimization, with the complete path intact and the fallback declared.\n\n**Every scar ends in a test.** Each real incident leaves its round with an automated check, not just a paragraph. A scar without a test reopens.\n\n**Reproducibility.** Same profile + same card = same result. Debuggable and testable like a factory part, not like a conversational black box.\n\n**Model per task.** A cheap model for the dumb work, an expensive one only for the hard part. A single agent uses one model for everything.\n\n**Visible cost, per piece.** Each closed card carries its model and its cost. You see spend per task and per flow, not an opaque monthly bill.\n\n**No provider lock-in.** No model is hardcoded anywhere. Switching providers is editing a table. In this field the landscape changes every quarter.\n\n**Horizontal scale.** Several machines can share one board. (If these were concurrent sub-second tasks in a critical system you'd want a real database. For chunked agentic work, a folder of files is plenty.)\n\n**Parallelism without collisions.** Many workers at once, each isolated. Claiming is moving a file, so two never collide.\n\n**Auditability with no extra tooling.** State IS the board, and a closed card is the receipt. All plain text, no external dashboard to maintain.\n\n**Growth costs no friction.** Adding an agent requires modifying nothing, because there's nobody to notify: integration happens by dropping files on the board.\n\n**Pull, not push.** Nothing is produced without a signal to pull. Idle capacity waiting beats accumulated inventory. And you stop being the central planner who matches every task to every tool by hand.\n\n**If the tool disappears, you keep the method.** The whole system is a folder of text files. It's exactly what Toyota would keep if you took away their machines: the method, not the equipment.\n\nIt is not a framework, a library, or a product. There is nothing to `pip install`\n\n.\n\nIt is not a replacement for conversational agents. Delegating to a subagent inside a session you're actively running is the right call for heavy work you need in the next two minutes. This standard is for work you want to happen without you, repeatably, and be shareable.\n\nIt is not novel infrastructure. Dependency graphs are from the seventies, `make`\n\nis from 1976, and kanban is from Toyota in the fifties. What's new here is the substrate: applying it to LLM agents defined as plain text contracts, so they survive the tool that runs them.\n\nBuilt by **Marcos Emowe**, 2026. Running in production on a Mac Mini since August 2026. I write about digital brains and AI agents at [emowe.com](https://emowe.com).\n\nThe reasoning behind this standard, in Spanish: [por qué atomizamos](https://emowe.com/cerebro-digital/por-que-atomizamos/).\n\nIf you build on this, a link back is appreciated. The specification is free to read, use and adapt. The name **Cerebro Digital®** is a registered trademark; using it requires permission.", "url": "https://wpnews.pro/news/why-i-stopped-building-ai-agents-as-conversational-sessions-and-started-building", "canonical_source": "https://gist.github.com/emowe/b3563cbebaf060c998af829ab814dc30", "published_at": "2026-08-27 10:50:14+00:00", "updated_at": "2026-08-31 00:23:00.088395+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "ai-infrastructure"], "entities": ["Marcos Emowe", "Atomic Agents"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/why-i-stopped-building-ai-agents-as-conversational-sessions-and-started-building", "markdown": "https://wpnews.pro/news/why-i-stopped-building-ai-agents-as-conversational-sessions-and-started-building.md", "text": "https://wpnews.pro/news/why-i-stopped-building-ai-agents-as-conversational-sessions-and-started-building.txt", "jsonld": "https://wpnews.pro/news/why-i-stopped-building-ai-agents-as-conversational-sessions-and-started-building.jsonld"}}