My VPS runs a "personal AI agent". It forgets its own abilities every morning.
My terminal runs a coding agent. It ships production work.
Same year. Frontier models on both. Same ecosystem.
Both are model + harness, trying to handle the same thing: my intent. Why such a
different experience? Start with the thing everyone mixes up: definitions.
Providers are wrapping yesterday's chats in agent loops. Execution pattern
flips, the chat UI stays. Same split still holds: model caps capability, harness
wires integrations and workflow, and intent has to be decomposed into pieces the
agent can handle. Who does the decomposition is the next question.
First, split common LLM application architectures by workflow: the execution
pattern around model calls.
At tiers 0-3 this is mostly application code. Fixed calls, branching, one-off
tool use. At tiers 4-5 it becomes an agent harness. Loop, state, permissions,
memory, orchestration.
Six tiers, in order of escalating cost/complexity:
| Tier | Pattern | When | Cost shape |
|---|---|---|---|
| 0 | Single prompt | Text in, text out | 1 LLM call |
| 1 | Prompt chain | Multi-step but predictable pipeline | N LLM calls |
| 2 | Routing | Input-type dispatch into one of K branches | Router call + selected branch |
| 3 | Tool-use, single round | Need external data, one tool round suffices | Tool schema tokens + tool call + synthesis |
| 4 | Agent, ReAct loop | Iterative, unknown depth, state-dependent | Unbounded: loop iterations × model/tool calls |
| 5 | Multi-agent / orchestrator | Long-horizon, decomposable, large context | Orchestrator + workers + replans + synthesis |
This post is about tiers 4 and 5. Agent loops, orchestrators, and who owns
intent decomposition.
Three families compete for "AI agent" today. Same lens for each: who owns intent
decomposition? The specialized harness, the personal assistant, or me.
Same product shape. Narrow task, fixed workflow, integrations hidden behind the
UI.
Perplexity is the research/search version of that shape.
Research the latest evidence and debate around emergent agency in AI systems
in 2025-2026.
-
Searching the web
-
... some sources ...
-
Searching recent academic and policy sources on emergent agency in AI systems
-
Retrieving recent reports and policy documents on AI agent behavior and emergent agency
Fine results, at least looks relevant.
what's the model you using?
I'm Perplexity, powered by Grok 4.1. Perplexity's default model is a
custom-optimized version called Sonar, built on Llama 3.1 or later
iterations like Llama 3.3 70B, fine-tuned for fast web search,
summarization, and efficiency. It's available to both free and Pro users for
quick, relevant results. Pro subscribers can switch to advanced options like
GPT-5.2 or Claude models via the model selector.
Fail. It understood the surface intent (answer a product-stack question), but
the product owned routing and context. I had no way to inspect which model
routed the answer, which metadata it saw, or why it mixed Grok, Sonar, and Llama
into one pile.
The outcome: specialized harness frames intent into its fixed shape. When that
frame fits, I get a clean research answer. When the frame itself is wrong, I get
confident product salad and no useful control surface.
Tried OpenClaw first. It wants 2+ CPUs and 8+ GB RAM; my VPS has 1 and 1. Ran
it anyway. It choked the VPS. Dropped it for Hermes.
Rarely discussed, but experimental software with a lot of external integrations
has too broad an attack surface, see https://days-since-openclaw-cve.com. Keep
it in mind.
Strange that Nous Research doesn't mention they have a Docker image,
docker.io/nousresearch/hermes-agent, which I've successfully set up in
[Unit]
Description=Hermes Agent
Wants=network-online.target
After=network-online.target
[Container]
Image=docker.io/nousresearch/hermes-agent:latest
ContainerName=hermes-agent
Network=selfhosted
Volume=/root/hermes:/opt/data
Volume=/root/hermes-root:/root
Volume=/tmp/hermes:/tmp
Ulimit=nofile=1024:1024
Environment=VIRTUAL_ENV=/root/.venv
Environment=PYTHONPATH=/root/.venv/lib/python3.13/site-packages
Exec=gateway run
[Service]
Restart=always
RestartSec=3
MemoryMax=768M
MemorySwapMax=768M
CPUQuota=85%
TasksMax=128
[Install]
WantedBy=multi-user.target
And it runs completely fine on a 1 CPU / 1 GB VPS.
CPU/RAM consumption I connected it to my GPT subscription, added integrations for X, Google
Calendar, Notion, and this blog's RSS, plus free Mem0 as RAG.
It even worked right after setup, but the next day it forgot about the
integration. I had to persuade it to try again.
OAuth failed in a different way. During Google Calendar integration I issued
credentials only for read/write on the calendar, not broader Google scopes. The
builtin Google skill wants broader access, so the agent re-requests broader
scopes every time it touches the calendar, and eventually the auth flow breaks
again.
One more case: I configured a scheduled job to check, each morning at 9:00, my
Notion calendar, Google Calendar, event listings, and send me a summary for
today and tomorrow. How often does it work right? Almost never. It checks only
one calendar, sends events for the next ~6 months instead of 2 days, sends
events for the current month but from this and previous years, and so on.
Current state: the initial GPT auth token has expired, and the agent can't renew
it automatically. Well... experiment successful.
Each failure is easy to fix manually. Cron, small script, explicit OAuth scopes,
date windows, deterministic calendar queries. But that is exactly the point: the
general assistant is supposed to replace the glue. Here it doesn't. The failure
is not the model. The harness decomposes intent badly, and the UI doesn't expose
decomposition early enough to fix it.<sup>3</sup>
Terminal-native, actively evolving. Everything in my hands. Only vendor ToS can
limit me.
My favorite one. Open source, standard ~/.config/opencode path, strong
build/plan sub-agent architecture. Also ships opencode web, same engine,
browser UI.
{
"$schema": "https://opencode.ai/config.json",
"autoupdate": false,
"default_agent": "plan",
"share": "disabled",
"snapshot": false,
"instructions": ["/Users/ivan/.config/opencode/AGENTS.md"],
"mcp": {
"context7": {
"type": "remote",
"url": "https://mcp.context7.com/mcp",
"headers": {
"CONTEXT7_API_KEY": "{env:CONTEXT7_API_KEY}"
},
"enabled": true
},
"playwright": {
"type": "local",
"command": [
"npx",
"-y",
"@playwright/mcp@latest",
"--browser=chromium",
"--executable-path=/Applications/Chromium.app/Contents/MacOS/Chromium",
"--caps=vision,devtools"
],
"environment": {
"PLAYWRIGHT_BROWSERS_PATH": "{env:HOME}/.cache/ms-playwright"
},
"enabled": true
}
},
"experimental": {
"openTelemetry": false
},
"server": {
"mdns": false
},
"plugin": []
}
Current AI-bro consensus: "Context is the key". I agree. Context engineering
pwnd prompt engineering.
Even more, after
Gloaguen et al., "Evaluating AGENTS.md" (Feb 2026),
generated agent context files hurt task success and add ~20% inference cost. I
stopped keeping per-repo CLAUDE.md / AGENTS.md files full of paths, framework
summaries, and obvious project descriptions. Current agents can inspect a repo
per case, fast and cheap enough. Put only what they cannot infer from code.
Constraints, preferences, dangerous commands, external contracts.
Be brief
Truth over comfort
Simple over clever
Contradiction: name both sides, never average
Use subagents to do the work; main thread is orchestrator
grep -> rg; python -> python3
Still, the most reliable way to fix hallucinations is not to argue with the
agent but to drop the session and restart from scratch.
OpenCode's approach helps me clearly follow the principles above:
In plan mode, the harness exposes decomposition before execution. At that point,
the limit is me: how clearly I can split intent into steps.
This resembles the Plan-Then-Execute pattern<sup>6</sup>.
I migrated Claude work to Claude Code after Anthropic restricted third-party
Claude access to API credits<sup>7</sup>; OpenCode still runs my GPT subscription.
Proprietary, vendor-locked, but I can't complain that it misses anything
important.
{
"$schema": "https://json.schemastore.org/claude-code-settings.json",
"cleanupPeriodDays": 1,
"env": {
"CLAUDE_CODE_DISABLE_FEEDBACK_SURVEY": "1",
"CLAUDE_CODE_DISABLE_TERMINAL_TITLE": "1",
"DISABLE_AUTOUPDATER": "1",
"DISABLE_BUG_COMMAND": "1",
"DISABLE_COST_WARNINGS": "1",
"CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC": "1",
"DISABLE_ERROR_REPORTING": "1",
"DISABLE_FEEDBACK_COMMAND": "1",
"DISABLE_TELEMETRY": "1"
},
"permissions": {
"allow": [
"Bash(git log *)",
"Bash(git diff *)",
"Bash(git show *)",
"Bash(git status)",
"Bash(grep *)",
"Bash(echo *)",
"Bash(ls *)",
"Bash(rg *)",
"Bash(npm run *)",
"Bash(npm test *)"
],
"deny": [
"Bash(curl *)",
"Bash(docker push *)",
"Bash(find * -delete)",
"Bash(find * -exec rm*)",
"Bash(git branch -D *)",
"Bash(git checkout -- *)",
"Bash(git clean -f*)",
"Bash(git push *)",
"Bash(git reset --hard*)",
"Bash(nc *)",
"Bash(rm -f *)",
"Bash(rm -r *)",
"Bash(rm -rf *)",
"Bash(rsync *)",
"Bash(scp *)",
"Bash(ssh *)",
"Bash(sudo *)",
"Bash(wget *)",
"Edit(./.env*)",
"Edit(./.git/**)",
"Edit(./secrets/**)",
"Edit(~/.aws/**)",
"Edit(~/.bashrc)",
"Edit(~/.ssh/**)",
"Edit(~/.zshrc)",
"Read(*.env)",
"Read(./.env.*)",
"Read(./secrets/**)",
"Read(~/.aws/**)",
"Read(~/.azure/**)",
"Read(~/.config/gh/**)",
"Read(~/.git-credentials)",
"Read(~/.gnupg/**)",
"Read(~/.kube/**)",
"Read(~/.npmrc)",
"Read(~/.ssh/**)"
],
"defaultMode": "plan"
},
"model": "sonnet",
"disableClaudeAiConnectors": true,
"disableBundledSkills": true,
"disableRemoteControl": true,
"disableWorkflows": true,
"disableArtifact": true,
"enableAllProjectMcpServers": false,
"includeCoAuthoredBy": false,
"sandbox": {
"enabled": true,
"excludedCommands": ["git"]
},
"effortLevel": "low",
"awaySummaryEnabled": false,
"autoUpdatesChannel": "stable",
"autoMemoryEnabled": false,
"disableAutoMode": "disable",
"theme": "light",
"editorMode": "normal",
"preferredNotifChannel": "notifications_disabled",
"autoCompactEnabled": false,
"skipAutoPermissionPrompt": false
}
This config reaches the same control goal I like in OpenCode: decomposition
stays visible, writes stay gated, mode switching stays under my control:
"defaultMode": "plan": plan mode by default. Writes blocked, reads allowed,
plan exposed as /plan
"disableAutoMode": "disable": no autonomous mode switching"sandbox.enabled": true plus a permissions.deny list for rm -rf,
git push, sudo, ~/.ssh/**, etc. DISABLE_TELEMETRY=1, DISABLE_FEEDBACK_COMMAND=1,
DISABLE_ERROR_REPORTING=1, awaySummaryEnabled: false
Each of these touches the harness only. Model stays the vendor's, intent stays
mine.
Perplexity frames intent into its fixed shape. Hermes decomposes intent on its
own, without me. CLI plan mode keeps decomposition visible. I'm the bottleneck.
The loop is what makes something agentic, and the harness puts me inside or
outside of it.
Same year. Same frontier models. Humans still hold the loop. That is the
autonomy that emerges today.<sup>10</sup>
{data-content="footnotes"}
Popularized by ReAct: Yao et al., ICLR 2023
(paper,
See Birgitta Böckeler, "Harness engineering for coding agent users":
https://martinfowler.com/articles/harness-engineering.html. ↩
Personal assistants still look worth watching. I'll wait for the next
Hermes iteration. Notion's agent may become the boring alternative. ↩
Yes-yes, I know about openspec.dev, but the plan has to stay observable, not
5+ A4 neuro-generated pages of
raw text. ↩
To save some tokens I use stronger model for plan (Opus) and weaker for
build (Sonnet). ↩
Caveat: the split is weaker here. Plan sub-agent still reads untrusted repo
while planning, so a malicious file can steer the plan. ↩
tightened terms. Subscription plans (including the corporate one I use) no
longer cover Claude access from third-party apps; third-party access now
requires per-usage API credits. ↩
CLI is more token effective than MCP,
https://github.com/microsoft/playwright-cli#playwright-cli-vs-playwright-mcp. ↩
Caveman vs "be brief",
https://www.maxtaylor.me/articles/i-benchmarked-caveman-against-two-words.
Agent tools move too fast; without a fresh benchmark after model or harness
updates, a third-party skill can quietly make results worse than the
baseline. Benchmark it continuously or keep it off. ↩
framed as emergent from model behavior, product design, and user oversight
strategy. ↩