AI agent architecture: model, harness and intent A developer's hands-on comparison of personal AI agents found that the same frontier models produce very different results depending on the harness around them, with a coding agent shipping production work while a VPS-hosted personal agent repeatedly forgot its own capabilities. The writeup organizes LLM application architectures into six tiers of escalating cost and complexity, from single prompts through ReAct loops to multi-agent orchestrators, and argues that the specialized harness — not the model — owns intent decomposition. Testing Perplexity, OpenClaw and Hermes, the developer reports the product's routing and context were opaque, OpenClaw exceeded the 1 CPU / 1 GB VPS and was dropped, and Hermes was run via its Docker image under podman. My VPS runs a "personal AI agent". It forgets its own abilities every morning. My terminal runs a coding agent. It ships production work. Same year. Frontier models on both. Same ecosystem. Both are model + harness, trying to handle the same thing: my intent. Why such a different experience? Start with the thing everyone mixes up: definitions. Providers are wrapping yesterday's chats in agent loops. Execution pattern flips, the chat UI stays. Same split still holds: model caps capability, harness wires integrations and workflow, and intent has to be decomposed into pieces the agent can handle. Who does the decomposition is the next question. First, split common LLM application architectures by workflow: the execution pattern around model calls. At tiers 0-3 this is mostly application code. Fixed calls, branching, one-off tool use. At tiers 4-5 it becomes an agent harness. Loop, state, permissions, memory, orchestration. Six tiers, in order of escalating cost/complexity: | Tier | Pattern | When | Cost shape | |---|---|---|---| | 0 | Single prompt | Text in, text out | 1 LLM call | | 1 | Prompt chain | Multi-step but predictable pipeline | N LLM calls | | 2 | Routing | Input-type dispatch into one of K branches | Router call + selected branch | | 3 | Tool-use, single round | Need external data, one tool round suffices | Tool schema tokens + tool call + synthesis | | 4 | Agent, ReAct loop | Iterative, unknown depth, state-dependent | Unbounded: loop iterations × model/tool calls | | 5 | Multi-agent / orchestrator | Long-horizon, decomposable, large context | Orchestrator + workers + replans + synthesis | This post is about tiers 4 and 5. Agent loops, orchestrators, and who owns intent decomposition. Three families compete for "AI agent" today. Same lens for each: who owns intent decomposition? The specialized harness, the personal assistant, or me. Same product shape. Narrow task, fixed workflow, integrations hidden behind the UI. Perplexity is the research/search version of that shape. Research the latest evidence and debate around emergent agency in AI systems in 2025-2026. - Searching the web - ... some sources ... - Searching recent academic and policy sources on emergent agency in AI systems - Retrieving recent reports and policy documents on AI agent behavior and emergent agency Fine results, at least looks relevant. what's the model you using? I'm Perplexity, powered by Grok 4.1. Perplexity's default model is a custom-optimized version called Sonar, built on Llama 3.1 or later iterations like Llama 3.3 70B, fine-tuned for fast web search, summarization, and efficiency. It's available to both free and Pro users for quick, relevant results. Pro subscribers can switch to advanced options like GPT-5.2 or Claude models via the model selector. Fail. It understood the surface intent answer a product-stack question , but the product owned routing and context. I had no way to inspect which model routed the answer, which metadata it saw, or why it mixed Grok, Sonar, and Llama into one pile. The outcome: specialized harness frames intent into its fixed shape. When that frame fits, I get a clean research answer. When the frame itself is wrong, I get confident product salad and no useful control surface. Tried OpenClaw first. It wants 2+ CPUs and 8+ GB RAM; my VPS has 1 and 1. Ran it anyway. It choked the VPS. Dropped it for Hermes . Rarely discussed, but experimental software with a lot of external integrations has too broad an attack surface, see https://days-since-openclaw-cve.com https://days-since-openclaw-cve.com . Keep it in mind. Strange that Nous Research doesn't mention they have a Docker image, docker.io/nousresearch/hermes-agent , which I've successfully set up in podman https://bogomolov.work/blog/posts/the-actual-state-of-self-hosting-on-a-vps/ . Unit Description=Hermes Agent Wants=network-online.target After=network-online.target Container Image=docker.io/nousresearch/hermes-agent:latest ContainerName=hermes-agent Network=selfhosted Volume=/root/hermes:/opt/data Volume=/root/hermes-root:/root Volume=/tmp/hermes:/tmp Ulimit=nofile=1024:1024 Environment=VIRTUAL ENV=/root/.venv Environment=PYTHONPATH=/root/.venv/lib/python3.13/site-packages Exec=gateway run Service Restart=always RestartSec=3 MemoryMax=768M MemorySwapMax=768M CPUQuota=85% TasksMax=128 Install WantedBy=multi-user.target And it runs completely fine on a 1 CPU / 1 GB VPS. CPU/RAM consumption I connected it to my GPT subscription, added integrations for X, Google Calendar, Notion, and this blog's RSS, plus free Mem0 as RAG. It even worked right after setup, but the next day it forgot about the integration. I had to persuade it to try again. OAuth failed in a different way. During Google Calendar integration I issued credentials only for read/write on the calendar, not broader Google scopes. The builtin Google skill wants broader access, so the agent re-requests broader scopes every time it touches the calendar, and eventually the auth flow breaks again. One more case: I configured a scheduled job to check, each morning at 9:00, my Notion calendar, Google Calendar, event listings, and send me a summary for today and tomorrow. How often does it work right? Almost never. It checks only one calendar, sends events for the next ~6 months instead of 2 days, sends events for the current month but from this and previous years, and so on. Current state: the initial GPT auth token has expired, and the agent can't renew it automatically. Well... experiment successful. Each failure is easy to fix manually. Cron, small script, explicit OAuth scopes, date windows, deterministic calendar queries. But that is exactly the point: the general assistant is supposed to replace the glue. Here it doesn't. The failure is not the model. The harness decomposes intent badly, and the UI doesn't expose decomposition early enough to fix it.