{"slug": "ai-agent-architecture-model-harness-and-intent", "title": "AI agent architecture: model, harness and intent", "summary": "A developer's hands-on comparison of personal AI agents found that the same frontier models produce very different results depending on the harness around them, with a coding agent shipping production work while a VPS-hosted personal agent repeatedly forgot its own capabilities. The writeup organizes LLM application architectures into six tiers of escalating cost and complexity, from single prompts through ReAct loops to multi-agent orchestrators, and argues that the specialized harness — not the model — owns intent decomposition. Testing Perplexity, OpenClaw and Hermes, the developer reports the product's routing and context were opaque, OpenClaw exceeded the 1 CPU / 1 GB VPS and was dropped, and Hermes was run via its Docker image under podman.", "body_md": "My VPS runs a \"personal AI agent\". It forgets its own abilities every morning.\n\nMy terminal runs a coding agent. It ships production work.\n\nSame year. Frontier models on both. Same ecosystem.\n\nBoth are model + harness, trying to handle the same thing: my intent. Why such a\n\ndifferent experience? Start with the thing everyone mixes up: definitions.\n\nProviders are wrapping yesterday's chats in agent loops. Execution pattern\n\nflips, the chat UI stays. Same split still holds: model caps capability, harness\n\nwires integrations and workflow, and intent has to be decomposed into pieces the\n\nagent can handle. Who does the decomposition is the next question.\n\nFirst, split common LLM application architectures by workflow: the execution\n\npattern around model calls.\n\nAt tiers 0-3 this is mostly application code. Fixed calls, branching, one-off\n\ntool use. At tiers 4-5 it becomes an agent harness. Loop, state, permissions,\n\nmemory, orchestration.\n\nSix tiers, in order of escalating cost/complexity:\n\n| Tier | Pattern | When | Cost shape | \n|---|---|---|---|\n| 0 | Single prompt | Text in, text out | 1 LLM call | \n| 1 | Prompt chain | Multi-step but predictable pipeline | N LLM calls | \n| 2 | Routing | Input-type dispatch into one of K branches | Router call + selected branch | \n| 3 | Tool-use, single round | Need external data, one tool round suffices | Tool schema tokens + tool call + synthesis | \n| 4 | Agent, ReAct loop | Iterative, unknown depth, state-dependent | Unbounded: loop iterations × model/tool calls | \n| 5 | Multi-agent / orchestrator | Long-horizon, decomposable, large context | Orchestrator + workers + replans + synthesis | \n\nThis post is about tiers 4 and 5. Agent loops, orchestrators, and who owns\n\nintent decomposition.\n\nThree families compete for \"AI agent\" today. Same lens for each: who owns intent\n\ndecomposition? The specialized harness, the personal assistant, or me.\n\nSame product shape. Narrow task, fixed workflow, integrations hidden behind the\n\nUI.\n\n*Perplexity* is the research/search version of that shape.\n\nResearch the latest evidence and debate around emergent agency in AI systems\n\nin 2025-2026.\n\n- Searching the web\n\n- ... some sources ...\n- Searching recent academic and policy sources on emergent agency in AI systems\n- Retrieving recent reports and policy documents on AI agent behavior and emergent agency\n\nFine results, at least looks relevant.\n\nwhat's the model you using?\n\nI'm Perplexity, powered by Grok 4.1. Perplexity's default model is a\n\ncustom-optimized version called Sonar, built on Llama 3.1 or later\n\niterations like Llama 3.3 70B, fine-tuned for fast web search,\n\nsummarization, and efficiency. It's available to both free and Pro users for\n\nquick, relevant results. Pro subscribers can switch to advanced options like\n\nGPT-5.2 or Claude models via the model selector.\n\nFail. It understood the surface intent (answer a product-stack question), but\n\nthe product owned routing and context. I had no way to inspect which model\n\nrouted the answer, which metadata it saw, or why it mixed Grok, Sonar, and Llama\n\ninto one pile.\n\nThe outcome: specialized harness frames intent into its fixed shape. When that\n\nframe fits, I get a clean research answer. When the frame itself is wrong, I get\n\nconfident product salad and no useful control surface.\n\nTried *OpenClaw* first. It wants 2+ CPUs and 8+ GB RAM; my VPS has 1 and 1. Ran\n\nit anyway. It choked the VPS. Dropped it for *Hermes*.\n\nRarely discussed, but experimental software with a lot of external integrations\n\nhas too broad an attack surface, see [https://days-since-openclaw-cve.com](https://days-since-openclaw-cve.com). Keep\n\nit in mind.\n\nStrange that Nous Research doesn't mention they have a Docker image,\n\n`docker.io/nousresearch/hermes-agent`, which I've successfully set up in\n\n[podman](https://bogomolov.work/blog/posts/the-actual-state-of-self-hosting-on-a-vps/).\n\n```\n[Unit]\nDescription=Hermes Agent\nWants=network-online.target\nAfter=network-online.target\n\n[Container]\nImage=docker.io/nousresearch/hermes-agent:latest\nContainerName=hermes-agent\nNetwork=selfhosted\nVolume=/root/hermes:/opt/data\nVolume=/root/hermes-root:/root\nVolume=/tmp/hermes:/tmp\nUlimit=nofile=1024:1024\n\nEnvironment=VIRTUAL_ENV=/root/.venv\nEnvironment=PYTHONPATH=/root/.venv/lib/python3.13/site-packages\n\nExec=gateway run\n\n[Service]\nRestart=always\nRestartSec=3\nMemoryMax=768M\nMemorySwapMax=768M\nCPUQuota=85%\nTasksMax=128\n\n[Install]\nWantedBy=multi-user.target\n```\n\nAnd it runs completely fine on a 1 CPU / 1 GB VPS.\n\nCPU/RAM consumption\nI connected it to my GPT subscription, added integrations for X, Google\n\nCalendar, Notion, and this blog's RSS, plus free Mem0 as RAG.\n\nIt even worked right after setup, but the next day it forgot about the\n\nintegration. I had to persuade it to try again.\n\nOAuth failed in a different way. During Google Calendar integration I issued\n\ncredentials only for read/write on the calendar, not broader Google scopes. The\n\nbuiltin Google skill wants broader access, so the agent re-requests broader\n\nscopes every time it touches the calendar, and eventually the auth flow breaks\n\nagain.\n\nOne more case: I configured a scheduled job to check, each morning at 9:00, my\n\nNotion calendar, Google Calendar, event listings, and send me a summary for\n\ntoday and tomorrow. How often does it work right? Almost never. It checks only\n\none calendar, sends events for the next ~6 months instead of 2 days, sends\n\nevents for the current month but from this and previous years, and so on.\n\nCurrent state: the initial GPT auth token has expired, and the agent can't renew\n\nit automatically. Well... experiment successful.\n\nEach failure is easy to fix manually. Cron, small script, explicit OAuth scopes,\n\ndate windows, deterministic calendar queries. But that is exactly the point: the\n\ngeneral assistant is supposed to replace the glue. Here it doesn't. The failure\n\nis not the model. The harness decomposes intent badly, and the UI doesn't expose\n\ndecomposition early enough to fix it.<sup>3</sup>\n\nTerminal-native, actively evolving. Everything in my hands. Only vendor ToS can\n\nlimit me.\n\nMy favorite one. Open source, standard `~/.config/opencode` path, strong\n\nbuild/plan sub-agent architecture. Also ships `opencode web`, same engine,\n\nbrowser UI.\n\n```\n{\n  \"$schema\": \"https://opencode.ai/config.json\",\n  \"autoupdate\": false,\n  \"default_agent\": \"plan\",\n  \"share\": \"disabled\",\n  \"snapshot\": false,\n  \"instructions\": [\"/Users/ivan/.config/opencode/AGENTS.md\"],\n  \"mcp\": {\n    \"context7\": {\n      \"type\": \"remote\",\n      \"url\": \"https://mcp.context7.com/mcp\",\n      \"headers\": {\n        \"CONTEXT7_API_KEY\": \"{env:CONTEXT7_API_KEY}\"\n      },\n      \"enabled\": true\n    },\n    \"playwright\": {\n      \"type\": \"local\",\n      \"command\": [\n        \"npx\",\n        \"-y\",\n        \"@playwright/mcp@latest\",\n        \"--browser=chromium\",\n        \"--executable-path=/Applications/Chromium.app/Contents/MacOS/Chromium\",\n        \"--caps=vision,devtools\"\n      ],\n      \"environment\": {\n        \"PLAYWRIGHT_BROWSERS_PATH\": \"{env:HOME}/.cache/ms-playwright\"\n      },\n      \"enabled\": true\n    }\n  },\n  \"experimental\": {\n    \"openTelemetry\": false\n  },\n  \"server\": {\n    \"mdns\": false\n  },\n  \"plugin\": []\n}\n```\n\nCurrent AI-bro consensus: \"Context is the key\". I agree. Context engineering\n\npwnd [prompt engineering](https://bogomolov.work/blog/posts/prompt-engineering-notes/).\n\nEven more, after\n\n[Gloaguen et al., \"Evaluating AGENTS.md\" (Feb 2026)](https://arxiv.org/abs/2602.11988),\n\ngenerated agent context files hurt task success and add ~20% inference cost. I\n\nstopped keeping per-repo CLAUDE.md / AGENTS.md files full of paths, framework\n\nsummaries, and obvious project descriptions. Current agents can inspect a repo\n\nper case, fast and cheap enough. Put only what they cannot infer from code.\n\nConstraints, preferences, dangerous commands, external contracts.\n\n```\n# Invariants\n\nBe brief\n\nTruth over comfort\n\nSimple over clever\n\nContradiction: name both sides, never average\n\nUse subagents to do the work; main thread is orchestrator\n\ngrep -> rg; python -> python3\n```\n\nStill, the most reliable way to fix hallucinations is not to argue with the\n\nagent but to drop the session and restart from scratch.\n\nOpenCode's approach helps me clearly follow the principles above:\n\nIn plan mode, the harness exposes decomposition before execution. At that point,\n\nthe limit is me: how clearly I can split intent into steps.\n\nThis resembles the Plan-Then-Execute pattern<sup>6</sup>.\n\nI migrated Claude work to Claude Code after Anthropic restricted third-party\n\nClaude access to API credits<sup>7</sup>; OpenCode still runs my GPT subscription.\n\nProprietary, vendor-locked, but I can't complain that it misses anything\n\nimportant.\n\n```\n{\n  \"$schema\": \"https://json.schemastore.org/claude-code-settings.json\",\n  \"cleanupPeriodDays\": 1,\n  \"env\": {\n    \"CLAUDE_CODE_DISABLE_FEEDBACK_SURVEY\": \"1\",\n    \"CLAUDE_CODE_DISABLE_TERMINAL_TITLE\": \"1\",\n    \"DISABLE_AUTOUPDATER\": \"1\",\n    \"DISABLE_BUG_COMMAND\": \"1\",\n    \"DISABLE_COST_WARNINGS\": \"1\",\n    \"CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC\": \"1\",\n    \"DISABLE_ERROR_REPORTING\": \"1\",\n    \"DISABLE_FEEDBACK_COMMAND\": \"1\",\n    \"DISABLE_TELEMETRY\": \"1\"\n  },\n  \"permissions\": {\n    \"allow\": [\n      \"Bash(git log *)\",\n      \"Bash(git diff *)\",\n      \"Bash(git show *)\",\n      \"Bash(git status)\",\n      \"Bash(grep *)\",\n      \"Bash(echo *)\",\n      \"Bash(ls *)\",\n      \"Bash(rg *)\",\n      \"Bash(npm run *)\",\n      \"Bash(npm test *)\"\n    ],\n    \"deny\": [\n      \"Bash(curl *)\",\n      \"Bash(docker push *)\",\n      \"Bash(find * -delete)\",\n      \"Bash(find * -exec rm*)\",\n      \"Bash(git branch -D *)\",\n      \"Bash(git checkout -- *)\",\n      \"Bash(git clean -f*)\",\n      \"Bash(git push *)\",\n      \"Bash(git reset --hard*)\",\n      \"Bash(nc *)\",\n      \"Bash(rm -f *)\",\n      \"Bash(rm -r *)\",\n      \"Bash(rm -rf *)\",\n      \"Bash(rsync *)\",\n      \"Bash(scp *)\",\n      \"Bash(ssh *)\",\n      \"Bash(sudo *)\",\n      \"Bash(wget *)\",\n      \"Edit(./.env*)\",\n      \"Edit(./.git/**)\",\n      \"Edit(./secrets/**)\",\n      \"Edit(~/.aws/**)\",\n      \"Edit(~/.bashrc)\",\n      \"Edit(~/.ssh/**)\",\n      \"Edit(~/.zshrc)\",\n      \"Read(*.env)\",\n      \"Read(./.env.*)\",\n      \"Read(./secrets/**)\",\n      \"Read(~/.aws/**)\",\n      \"Read(~/.azure/**)\",\n      \"Read(~/.config/gh/**)\",\n      \"Read(~/.git-credentials)\",\n      \"Read(~/.gnupg/**)\",\n      \"Read(~/.kube/**)\",\n      \"Read(~/.npmrc)\",\n      \"Read(~/.ssh/**)\"\n    ],\n    \"defaultMode\": \"plan\"\n  },\n  \"model\": \"sonnet\",\n  \"disableClaudeAiConnectors\": true,\n  \"disableBundledSkills\": true,\n  \"disableRemoteControl\": true,\n  \"disableWorkflows\": true,\n  \"disableArtifact\": true,\n  \"enableAllProjectMcpServers\": false,\n  \"includeCoAuthoredBy\": false,\n  \"sandbox\": {\n    \"enabled\": true,\n    \"excludedCommands\": [\"git\"]\n  },\n  \"effortLevel\": \"low\",\n  \"awaySummaryEnabled\": false,\n  \"autoUpdatesChannel\": \"stable\",\n  \"autoMemoryEnabled\": false,\n  \"disableAutoMode\": \"disable\",\n  \"theme\": \"light\",\n  \"editorMode\": \"normal\",\n  \"preferredNotifChannel\": \"notifications_disabled\",\n  \"autoCompactEnabled\": false,\n  \"skipAutoPermissionPrompt\": false\n}\n```\n\nThis config reaches the same control goal I like in OpenCode: decomposition\n\nstays visible, writes stay gated, mode switching stays under my control:\n\n`\"defaultMode\": \"plan\"`: plan mode by default. Writes blocked, reads allowed,\nplan exposed as `/plan`\n`\"disableAutoMode\": \"disable\"`: no autonomous mode switching`\"sandbox.enabled\": true` plus a `permissions.deny` list for `rm -rf`,\n`git push`, `sudo`, `~/.ssh/**`, etc.` DISABLE_TELEMETRY=1`, `DISABLE_FEEDBACK_COMMAND=1`,\n`DISABLE_ERROR_REPORTING=1`, `awaySummaryEnabled: false`\nEach of these touches the harness only. Model stays the vendor's, intent stays\n\nmine.\n\nPerplexity frames intent into its fixed shape. Hermes decomposes intent on its\n\nown, without me. CLI plan mode keeps decomposition visible. I'm the bottleneck.\n\n*The loop is what makes something agentic*, and the harness puts me inside or\n\noutside of it.\n\nSame year. Same frontier models. Humans still hold the loop. That is the\n\nautonomy that emerges today.<sup>10</sup>\n\n{data-content=\"footnotes\"}\n\nPopularized by ReAct: Yao et al., ICLR 2023\n\n([paper](https://arxiv.org/abs/2210.03629),\n\n[Google Research blog](https://research.google/blog/react-synergizing-reasoning-and-acting-in-language-models/)). ↩\n\nSee Birgitta Böckeler, \"Harness engineering for coding agent users\":\n\n[https://martinfowler.com/articles/harness-engineering.html](https://martinfowler.com/articles/harness-engineering.html). ↩\n\n*Personal assistants* still look worth watching. I'll wait for the next\n\nHermes iteration. Notion's agent may become the boring alternative. ↩\n\nYes-yes, I know about openspec.dev, but the plan has to stay observable, not\n\n5+ A4 [neuro-generated](https://bogomolov.work/blog/posts/rotten-specs/) pages of\n\nraw text. ↩\n\nTo save some tokens I use stronger model for plan (Opus) and weaker for\n\nbuild (Sonnet). ↩\n\n[https://simonwillison.net/2025/Jun/13/prompt-injection-design-patterns/#the-plan-then-execute-pattern](https://simonwillison.net/2025/Jun/13/prompt-injection-design-patterns/#the-plan-then-execute-pattern)\n\nCaveat: the split is weaker here. Plan sub-agent still reads untrusted repo\n\nwhile planning, so a malicious file can steer the plan. ↩\n\n[Anthropic](https://bogomolov.work/blog/posts/will-ai-replace-developers/)\n\ntightened terms. Subscription plans (including the corporate one I use) no\n\nlonger cover Claude access from third-party apps; third-party access now\n\nrequires per-usage API credits. ↩\n\nCLI is more token effective than MCP,\n\n[https://github.com/microsoft/playwright-cli#playwright-cli-vs-playwright-mcp](https://github.com/microsoft/playwright-cli#playwright-cli-vs-playwright-mcp). ↩\n\nCaveman vs \"be brief\",\n\n[https://www.maxtaylor.me/articles/i-benchmarked-caveman-against-two-words](https://www.maxtaylor.me/articles/i-benchmarked-caveman-against-two-words).\n\nAgent tools move too fast; without a fresh benchmark after model or harness\n\nupdates, a third-party skill can quietly make results worse than the\n\nbaseline. Benchmark it continuously or keep it off. ↩\n\n[Agent autonomy](https://www.anthropic.com/news/measuring-agent-autonomy)\n\nframed as emergent from model behavior, product design, and user oversight\n\nstrategy. ↩", "url": "https://wpnews.pro/news/ai-agent-architecture-model-harness-and-intent", "canonical_source": "https://dev.to/irr123456/ai-agent-architecture-model-harness-and-intent-3018", "published_at": "2026-10-03 23:31:16+00:00", "updated_at": "2026-10-03 23:37:51.372553+00:00", "lang": "en", "topics": ["ai-agents", "large-language-models", "ai-tools", "ai-infrastructure"], "entities": ["Perplexity", "OpenClaw", "Hermes", "Nous Research", "Grok 4.1", "Sonar", "Llama 3.3 70B", "podman"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/ai-agent-architecture-model-harness-and-intent", "markdown": "https://wpnews.pro/news/ai-agent-architecture-model-harness-and-intent.md", "text": "https://wpnews.pro/news/ai-agent-architecture-model-harness-and-intent.txt", "jsonld": "https://wpnews.pro/news/ai-agent-architecture-model-harness-and-intent.jsonld"}}