{"slug": "obelisk-0-42-durable-agents-layered-sandboxes", "title": "Obelisk 0.42: Durable Agents, Layered Sandboxes", "summary": "Obelisk released version 0.42, splitting its single server.toml policy file into three files with three owners — server.toml for the platform, app.toml for the app policy, and deployment.toml for the application — so that the effective permission is the intersection of the grants. The release adds native V8 as an opt-in JavaScript engine via OBELISK_JS_RUNTIME=v8 (Boa remains the default), Linux VM activities, and a workflow-agent prototype that can deploy, test, and fix applications on another Obelisk instance. Exec activities now require approval from both the platform admin in server.toml and the app admin in app.toml, and exposed secrets require a grant in app.toml bound to a digest of the component and its complete set of exposed secrets.", "body_md": "# Obelisk 0.42: Durable Agents, Layered Sandboxes\n\nAn agent's progress should survive the process that runs it. Its code should run within a policy you can review, and its failures should leave a history you can inspect.\n\nObelisk keeps workflow progress in a database. Workflow code replays deterministically from that history, using recorded activity results before doing new work. A process can stop; the next one reconstructs where it left off. The same history lets you see what ran and debug it afterward.\n\n0.42 builds on that foundation for agentic workloads: reviewable security boundaries for generated code, native V8 for faster JavaScript replay, Linux VM activities for tools that need them, and a workflow-agent prototype that can deploy, test, and fix applications on another Obelisk instance.\n\nThe idea follows\n[SQLite is All You Need for Durable Workflows](https://obeli.sk/blog/sqlite-is-all-you-need-for-durable-workflows/):\nkeep durable state close to the runtime, and let compute come and go. An agent waiting for a model,\na tool, or a person can be represented by rows in the database, without a running VM per session.\n\n## Security\n\n### Three files, three owners\n\nIn 0.41 the operator's policy lived in `server.toml` and the application lived in `deployment.toml`.\nThat left one file doing two jobs: `server.toml` described both the platform (listeners, database,\nresource limits) and what a particular app was allowed to do. 0.42 splits it:\n\nThe effective permission is the intersection. A deployment can request less than `app.toml` grants,\nnever more, and `app.toml` cannot switch on exec activities unless `server.toml` allows it.\n\nAn agent can rewrite code and `deployment.toml` within those grants. A request for broader access\nrequires a change to `app.toml`, which gives the reviewer a small, explicit policy diff: which\nsecrets the code can use, which hosts it can call, and which native executables it can run.\n\n```\n# app.toml\napp_name = \"my-app\"\n\n[secrets]\nOPENAI_KEY = {}\n\n[[outbound_http.allowed_host]]\npattern = \"api.openai.com\"\nmethods = [\"POST\"]\nrequest_url_regex = \"^POST https://api\\\\.openai\\\\.com/v1/\"\nsecrets = [\"OPENAI_KEY\"]\nreplace_in = [\"headers\"]\n```\n\n`obelisk deployment verify` reports missing policy entries; `--fix` can scaffold them for review.\nEvery deployment records the digest of the app policy it was activated under, so that boundary is\npart of its inspectable history.\n\n### Secrets stay at the network edge\n\nSecrets still default to placeholders that the runtime replaces at the network edge, so component\ncode never sees the value. Some code legitimately needs the plaintext, such as a webhook verifying\nan HMAC signature. In 0.42, WASM and JavaScript activities, webhooks, exec activities, and VM\nactivities can request that with `exposed_secrets`.\n\nEach exposure requires a grant in `app.toml` bound to a digest of the component and its complete set\nof exposed secrets. If the agent changes the component or asks for one more secret, the digest\nchanges and the grant no longer applies. The app admin must approve the new digest before the\nruntime exposes those secrets.\n\n### Exec activities require both admins' approval\n\nExisting exec activities run host processes outside the sandbox. In 0.42, both the platform admin\nand the app admin must approve them: the platform permits exec in `server.toml`, and the app grants\naccess in `app.toml`.\n\nThe [Security Model](https://obeli.sk/docs/v0.42.0/security/) explains the\ngrants and approval digests in detail.\n\n## Native V8\n\nJavaScript workflows, activities, and webhooks can now run on native V8 instead of Boa compiled to\nWASM: start the server with `OBELISK_JS_RUNTIME=v8`. Each activity gets a fresh isolate. Components\nusing the 0.42 JavaScript API need no changes to switch engines. Boa remains the default engine.\n\nConcurrency and memory are bounded per workload and runtime by `[limits]` in `server.toml`, so the\nplatform admin can cap V8 activities, WASM workflows, and VM activities independently.\n\nFaster replay means less time reconstructing a session before it can resume. Our durable\ncoding-agent prototype, [workflow-agent](https://obeli.sk/blog/announcing-obelisk-0-42/#agentic-workflows), has separate JavaScript and Rust\nworkflow implementations. We replayed the same real agent conversation with both, comparing V8, Boa\nWASM, and Rust in Wasmtime:\n\nFor this 726-event conversation, median replay fell from 3.17 seconds on Boa WASM to 128\nmilliseconds on V8, about 25× faster. Nine replays after warmup on the same Intel i9-14900HX host.\nBars show medians; labels show observed ranges. Replay time excludes database loading and varies by\nworkload. The [replay measurements](https://obeli.sk/blog/2026-10-04-obelisk-0-42/js-replay.json) include every\nsample and the benchmark setup.\n\n## VM activities (experimental)\n\nFor activities that need a real Linux userspace, 0.42 adds `[[activity_vm]]`: a script that runs\ninside a Linux VM, with its tools supplied as Nix store paths that are verified and mounted\nread-only.\n\nGuest HTTP goes through the same app and deployment policy as every other component, including secret placeholders, so a VM activity cannot reach a host that a JavaScript activity could not.\n\nChoose a backend by setting `OBELISK_UNSTABLE_ACTIVITY_VM` for both the CLI and the server:\n\n- `bochs-wasm` : the Bochs x86 emulator compiled to WASM, running inside Wasmtime. A Linux VM inside\nthe WASM sandbox, with no host binaries required. It is the slowest option and has a fixed 512 MiB\nguest.\n- `qemu-tcg` and`qemu-kvm` : native QEMU, with or without KVM, up to 16.25 GiB of guest RAM.\n- `firecracker` : a Firecracker microVM, cold booted for each execution; needs`/dev/kvm` .\n\nWhat does the VM layer cost? We measured from the activity's persisted `Locked` event to its\n`Finished` event using the published Obelisk 0.42.0 binary and published VM runtimes, including\nQEMU's 2026-10-01 EROFS bundles. The small cases print a string with Bash or use curl to fetch a\nlocal page through Obelisk's HTTP bridge. The larger case is the\n[inception Playwright demo](https://github.com/obeli-sk/demo-playwright/tree/main/inception): it\nlaunches Chromium, opens [trynix.dev](https://trynix.dev/) with Playwright, boots Obelisk in the\npage's Linux VM, runs `obelisk -v`, and returns the command output as the activity result. Its\ntiming includes that in-browser boot. Each bar is a median in seconds; the scales differ between\nworkloads.\n\nAll timings came from the same Intel i9-14900HX host. The VM runs were sequential, with cached\nruntime images and warmup runs. Bash printf and curl use 512 MiB and one guest vCPU; inception uses\n8 GiB and four. Bash and curl each have seven measured runs after two warmups; Chromium has three\nafter one warmup. The [benchmark notes](https://obeli.sk/blog/2026-10-04-obelisk-0-42/vm-benchmark/README.md)\ninclude the matching CSV measurements, commands, and pinned source and runtime versions.\n\nSee the\n[VM activity reference](https://obeli.sk/docs/v0.42.0/configuration/deployment/#experimental-vm-activities)\nfor configuration. The guest ABI, configuration, and behavior are experimental.\n\n## Agentic workflows\n\nThere are two ways to bring agentic workloads to Obelisk. Let a coding agent such as Claude Code or\nCodex generate application code, then deploy it within the application's security policy. Generated\napplications that do not call a model consume no further LLM tokens during execution. This follows\nthe idea in Kelsey Hightower's\n[Zero Token Architecture talk at PlatformCon 2026](https://www.youtube.com/watch?v=A7WFt2JQ5sg): use\nthe model to build the application, then run the resulting code.\n\nOr run the agent itself as a durable workflow, with model calls and tools as activities. This is useful for enterprise agents that wait on people or external systems, and for coding agents that work inside a simulated Bash session with a persistent virtual filesystem.\n\nEnterprise agents inherit parent/child agent hierarchies, durable scheduling, and pause/resume from\nthe runtime. Hierarchical cancellation requires every workflow on the cancellation path to be\nexplicitly marked with the `-cancellable` export suffix. Cancelled workflows do not run their own\ncleanup handlers; see\n[Structured Concurrency](https://obeli.sk/docs/v0.42.0/concepts/structured-concurrency/#cancellation)\nfor cleanup ownership. Obelisk can transparently unload inactive sessions and reconstruct them from\nrecorded history when work resumes. The same replay mechanism recovers their progress after a server\nrestart. These capabilities come with the workflow runtime.\n\n### A coding agent that can inspect what it ships\n\n[workflow-agent](https://github.com/obeli-sk/workflow-agent) is our prototype of that second path: a\nbrowser UI and a durable agent loop. For coding agents, its cheap virtual workspaces are just-bash\nsessions with a persistent virtual filesystem and deeply integrated Obelisk and MCP commands.\n\nIt also exposes a simulated `obelisk` CLI that can connect to a separate target Obelisk instance.\nThe agent can read the target's deployment into its virtual filesystem, edit application code, apply\nthe deployment, and test it by calling functions and webhooks. It can then inspect execution\nhistory, application logs, and recorded HTTP traces to diagnose failures, fix the code, and deploy\nagain. That introspection closes the loop between writing an app and checking how it actually runs.\n\nThe aim is thousands of concurrent sessions without an external VM per chat.\n\nThese agentic workflows also shape Obelisk's APIs. Earlier versions copied the growing conversation\ninto every LLM activity, ballooning workflow state and stored data. Now the workflow keeps the\nlatest reply, while the activity fetches previous messages in a batch. That prompted the new\n[batch events API](https://obeli.sk/docs/v0.42.0/access/api/#post-v1-executions-events-batch),\nwhich reads just the create and finish events of child executions. The\n[workflow-agent architecture](https://github.com/obeli-sk/workflow-agent#architecture) explains the\ndesign.\n\nFor a smaller starting point, [demo-agent](https://github.com/obeli-sk/demo-agent) provides a\nJavaScript agent loop, an LLM activity, example tools, a human-in-the-loop question, and a polling\nUI. Its mock deployment runs without an LLM key.\n\n## Web UI: themes and new screens\n\nThe Web UI gets light and dark themes, a refreshed layout, deployment graphs, and new screens for\nsystem events and retention. Application logs now have level and stream filters. It uses the REST\nAPI and ships in every 0.42 release binary, served at `http://localhost:8080` by default.\n\n## Also in 0.42\n\n- `obelisk generate new` creates a JavaScript starter app.\n- Retention and garbage collection run automatically (30 days by default) and can be managed with\n`obelisk admin` ; system events are persisted.\n- Compatible JavaScript and Rust workflow implementations can replay the same execution log, allowing a language switch mid-execution.\n- `execution submit --follow-logs` , SSE streams for`follow=true` , and a batch events endpoint.\n- Experimental WASIp3 support for WASM activities and webhooks.\n- gRPC and gRPC-web are deprecated; the REST `/v1` API covers everything they did.\n\n## Upgrading\n\nObelisk is no longer published to crates.io, so `cargo install obelisk` and `cargo binstall obelisk`\nno longer receive new versions. Choose a supported channel from the\n[installation guide](https://obeli.sk/install/).\n\nThis release breaks configuration, the JavaScript runtime API, WIT packages, and a few API\nendpoints. Split your configuration, name the app, import `obelisk:workflow@1.0.0` instead of using\nthe global `obelisk` object, rename `activity_exec.secrets` to `exposed_secrets` and generate the\ngrants, and review the new `[limits]` defaults. `deployment get` is now `deployment pull`.\n\nIf you use the default SQLite directory, pin the existing path or move the database before restarting.\n\nThe [Migrating to 0.42](https://obeli.sk/docs/v0.42.0/migrating-to-0.42/) guide\ncovers each step, and the\n[Security Model](https://obeli.sk/docs/v0.42.0/security/) page describes how the\nlayers fit together. The configuration reference is split the same way as the files:\n[server.toml](https://obeli.sk/docs/v0.42.0/configuration/server/),\n[app.toml](https://obeli.sk/docs/v0.42.0/configuration/app/), and\n[deployment.toml](https://obeli.sk/docs/v0.42.0/configuration/deployment/).\n\n## Full Changelog\n\nSee [CHANGELOG.md](https://github.com/obeli-sk/obelisk/blob/v0.42.0/CHANGELOG.md) for every change.", "url": "https://wpnews.pro/news/obelisk-0-42-durable-agents-layered-sandboxes", "canonical_source": "https://obeli.sk/blog/announcing-obelisk-0-42/", "published_at": "2026-10-04 14:30:55+00:00", "updated_at": "2026-10-04 14:43:05.878695+00:00", "lang": "en", "topics": ["ai-agents", "developer-tools", "ai-infrastructure", "ai-safety"], "entities": ["Obelisk", "V8", "Boa", "SQLite", "OpenAI"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/obelisk-0-42-durable-agents-layered-sandboxes", "markdown": "https://wpnews.pro/news/obelisk-0-42-durable-agents-layered-sandboxes.md", "text": "https://wpnews.pro/news/obelisk-0-42-durable-agents-layered-sandboxes.txt", "jsonld": "https://wpnews.pro/news/obelisk-0-42-durable-agents-layered-sandboxes.jsonld"}}