An agent's progress should survive the process that runs it. Its code should run within a policy you can review, and its failures should leave a history you can inspect.
Obelisk keeps workflow progress in a database. Workflow code replays deterministically from that history, using recorded activity results before doing new work. A process can stop; the next one reconstructs where it left off. The same history lets you see what ran and debug it afterward.
0.42 builds on that foundation for agentic workloads: reviewable security boundaries for generated code, native V8 for faster JavaScript replay, Linux VM activities for tools that need them, and a workflow-agent prototype that can deploy, test, and fix applications on another Obelisk instance.
The idea follows SQLite is All You Need for Durable Workflows: keep durable state close to the runtime, and let compute come and go. An agent waiting for a model, a tool, or a person can be represented by rows in the database, without a running VM per session.
Security #
Three files, three owners
In 0.41 the operator's policy lived in server.toml and the application lived in deployment.toml.
That left one file doing two jobs: server.toml described both the platform (listeners, database,
resource limits) and what a particular app was allowed to do. 0.42 splits it:
The effective permission is the intersection. A deployment can request less than app.toml grants,
never more, and app.toml cannot switch on exec activities unless server.toml allows it.
An agent can rewrite code and deployment.toml within those grants. A request for broader access
requires a change to app.toml, which gives the reviewer a small, explicit policy diff: which
secrets the code can use, which hosts it can call, and which native executables it can run.
app_name = "my-app"
[secrets]
OPENAI_KEY = {}
[[outbound_http.allowed_host]]
pattern = "api.openai.com"
methods = ["POST"]
request_url_regex = "^POST https://api\\.openai\\.com/v1/"
secrets = ["OPENAI_KEY"]
replace_in = ["headers"]
obelisk deployment verify reports missing policy entries; --fix can scaffold them for review.
Every deployment records the digest of the app policy it was activated under, so that boundary is
part of its inspectable history.
Secrets stay at the network edge
Secrets still default to placeholders that the runtime replaces at the network edge, so component
code never sees the value. Some code legitimately needs the plaintext, such as a webhook verifying
an HMAC signature. In 0.42, WASM and JavaScript activities, webhooks, exec activities, and VM
activities can request that with exposed_secrets.
Each exposure requires a grant in app.toml bound to a digest of the component and its complete set
of exposed secrets. If the agent changes the component or asks for one more secret, the digest
changes and the grant no longer applies. The app admin must approve the new digest before the
runtime exposes those secrets.
Exec activities require both admins' approval
Existing exec activities run host processes outside the sandbox. In 0.42, both the platform admin
and the app admin must approve them: the platform permits exec in server.toml, and the app grants
access in app.toml.
The Security Model explains the grants and approval digests in detail.
Native V8 #
JavaScript workflows, activities, and webhooks can now run on native V8 instead of Boa compiled to
WASM: start the server with OBELISK_JS_RUNTIME=v8. Each activity gets a fresh isolate. Components
using the 0.42 JavaScript API need no changes to switch engines. Boa remains the default engine.
Concurrency and memory are bounded per workload and runtime by [limits] in server.toml, so the
platform admin can cap V8 activities, WASM workflows, and VM activities independently.
Faster replay means less time reconstructing a session before it can resume. Our durable coding-agent prototype, workflow-agent, has separate JavaScript and Rust workflow implementations. We replayed the same real agent conversation with both, comparing V8, Boa WASM, and Rust in Wasmtime:
For this 726-event conversation, median replay fell from 3.17 seconds on Boa WASM to 128 milliseconds on V8, about 25× faster. Nine replays after warmup on the same Intel i9-14900HX host. Bars show medians; labels show observed ranges. Replay time excludes database and varies by workload. The replay measurements include every sample and the benchmark setup.
VM activities (experimental) #
For activities that need a real Linux userspace, 0.42 adds [[activity_vm]]: a script that runs
inside a Linux VM, with its tools supplied as Nix store paths that are verified and mounted
read-only.
Guest HTTP goes through the same app and deployment policy as every other component, including secret placeholders, so a VM activity cannot reach a host that a JavaScript activity could not.
Choose a backend by setting OBELISK_UNSTABLE_ACTIVITY_VM for both the CLI and the server:
bochs-wasm: the Bochs x86 emulator compiled to WASM, running inside Wasmtime. A Linux VM inside the WASM sandbox, with no host binaries required. It is the slowest option and has a fixed 512 MiB guest.qemu-tcgandqemu-kvm: native QEMU, with or without KVM, up to 16.25 GiB of guest RAM.firecracker: a Firecracker microVM, cold booted for each execution; needs/dev/kvm.
What does the VM layer cost? We measured from the activity's persisted Locked event to its
Finished event using the published Obelisk 0.42.0 binary and published VM runtimes, including
QEMU's 2026-10-01 EROFS bundles. The small cases print a string with Bash or use curl to fetch a
local page through Obelisk's HTTP bridge. The larger case is the
inception Playwright demo: it
launches Chromium, opens trynix.dev with Playwright, boots Obelisk in the
page's Linux VM, runs obelisk -v, and returns the command output as the activity result. Its
timing includes that in-browser boot. Each bar is a median in seconds; the scales differ between
workloads.
All timings came from the same Intel i9-14900HX host. The VM runs were sequential, with cached runtime images and warmup runs. Bash printf and curl use 512 MiB and one guest vCPU; inception uses 8 GiB and four. Bash and curl each have seven measured runs after two warmups; Chromium has three after one warmup. The benchmark notes include the matching CSV measurements, commands, and pinned source and runtime versions.
See the VM activity reference for configuration. The guest ABI, configuration, and behavior are experimental.
Agentic workflows #
There are two ways to bring agentic workloads to Obelisk. Let a coding agent such as Claude Code or Codex generate application code, then deploy it within the application's security policy. Generated applications that do not call a model consume no further LLM tokens during execution. This follows the idea in Kelsey Hightower's Zero Token Architecture talk at PlatformCon 2026: use the model to build the application, then run the resulting code.
Or run the agent itself as a durable workflow, with model calls and tools as activities. This is useful for enterprise agents that wait on people or external systems, and for coding agents that work inside a simulated Bash session with a persistent virtual filesystem.
Enterprise agents inherit parent/child agent hierarchies, durable scheduling, and /resume from
the runtime. Hierarchical cancellation requires every workflow on the cancellation path to be
explicitly marked with the -cancellable export suffix. Cancelled workflows do not run their own
cleanup handlers; see
Structured Concurrency
for cleanup ownership. Obelisk can transparently unload inactive sessions and reconstruct them from
recorded history when work resumes. The same replay mechanism recovers their progress after a server
restart. These capabilities come with the workflow runtime.
A coding agent that can inspect what it ships
workflow-agent is our prototype of that second path: a browser UI and a durable agent loop. For coding agents, its cheap virtual workspaces are just-bash sessions with a persistent virtual filesystem and deeply integrated Obelisk and MCP commands.
It also exposes a simulated obelisk CLI that can connect to a separate target Obelisk instance.
The agent can read the target's deployment into its virtual filesystem, edit application code, apply
the deployment, and test it by calling functions and webhooks. It can then inspect execution
history, application logs, and recorded HTTP traces to diagnose failures, fix the code, and deploy
again. That introspection closes the loop between writing an app and checking how it actually runs.
The aim is thousands of concurrent sessions without an external VM per chat.
These agentic workflows also shape Obelisk's APIs. Earlier versions copied the growing conversation into every LLM activity, ballooning workflow state and stored data. Now the workflow keeps the latest reply, while the activity fetches previous messages in a batch. That prompted the new batch events API, which reads just the create and finish events of child executions. The workflow-agent architecture explains the design.
For a smaller starting point, demo-agent provides a JavaScript agent loop, an LLM activity, example tools, a human-in-the-loop question, and a polling UI. Its mock deployment runs without an LLM key.
Web UI: themes and new screens #
The Web UI gets light and dark themes, a refreshed layout, deployment graphs, and new screens for
system events and retention. Application logs now have level and stream filters. It uses the REST
API and ships in every 0.42 release binary, served at http://localhost:8080 by default.
Also in 0.42 #
obelisk generate newcreates a JavaScript starter app.- Retention and garbage collection run automatically (30 days by default) and can be managed with
obelisk admin; system events are persisted. - Compatible JavaScript and Rust workflow implementations can replay the same execution log, allowing a language switch mid-execution.
execution submit --follow-logs, SSE streams forfollow=true, and a batch events endpoint.- Experimental WASIp3 support for WASM activities and webhooks.
- gRPC and gRPC-web are deprecated; the REST
/v1API covers everything they did.
Upgrading #
Obelisk is no longer published to crates.io, so cargo install obelisk and cargo binstall obelisk
no longer receive new versions. Choose a supported channel from the
installation guide.
This release breaks configuration, the JavaScript runtime API, WIT packages, and a few API
endpoints. Split your configuration, name the app, import obelisk:workflow@1.0.0 instead of using
the global obelisk object, rename activity_exec.secrets to exposed_secrets and generate the
grants, and review the new [limits] defaults. deployment get is now deployment pull.
If you use the default SQLite directory, pin the existing path or move the database before restarting.
The Migrating to 0.42 guide covers each step, and the Security Model page describes how the layers fit together. The configuration reference is split the same way as the files: server.toml, app.toml, and deployment.toml.
Full Changelog #
See CHANGELOG.md for every change.