{"slug": "leashterm-a-new-programming-language-for-agents", "title": "Leashterm – a new programming language for agents", "summary": "Leashterm, a new programming language for AI agents renamed from \"agentlang,\" enforces declared permissions, bounded retries, and a default 10,000-step budget as part of the language itself rather than through an after-the-fact sandbox. The project's v0.7 release adds a general step budget covered by 44 tests, while v0.6 passed its tests and a 22-task benchmark in Codespaces and on GitHub, including a reproducible 9-trial result. In case 1 (cases/case1-filesystem/), 3 of 9 trials did not attempt an undeclared read but still smuggled a dependency past Leashterm by writing source code that referenced something outside its permissions, which the author describes as the honest edge of a language-level boundary.", "body_md": "A tiny language for AI agents. Instead of letting a program do anything and bolting a sandbox on afterwards, the things agents need are part of the language. (Renamed from \"agentlang\", which turned out to already be the name of an unrelated, existing open-source project.)\n\n1. **Permissions are declared up front.** A program can only read or write what it\ndeclared with`needs` . Anything else is refused.\n2. **Retries are always bounded.**`retry N { ... }` needs a fixed`N` (1 to 10).\n3. **Verification is a statement.**`verify a == b` stops the program if it does not hold.\n4. **Every action is logged** in a hash-chained audit log that can be checked for tampering.\n5. **Errors are structured JSON** with a kind, a line, a message and a concrete hint,\nso a model can feed them back and repair its own program.\n6. **Every program terminates.** The only repetition is`retry` (fixed limit) and`for` over a finite list. There is no`while` .`if` /`else` only chooses between blocks.\n7. **Total work is bounded.** Every statement counts against a step budget\n(`--max-steps` , default 10,000), the general safety net that also stops nested`retry` blocks from silently multiplying their attempts.\n\n**Leashterm is less a general-purpose programming language and more an executable\ncapability manifest with computation attached.** `needs` says what a program can touch;\nthe absence of unbounded loops says how much computational escalation is possible; the\nstep budget bounds composite work; and the audit log makes executed behavior checkable\nafter the fact. The distinguishing core is not the syntax - it is pre-execution capability\nchecking plus structurally bounded computation, and that is the identity the language\nshould stay tightly built around as it grows.\n\nThe guiding design question for any future addition is not \"what features is the language\nmissing?\" but: **what is the smallest language in which an agent can still do useful work,\nwhile every program still admits a compact, pre-execution upper bound on both its\ncapabilities and its work?** That is also why arithmetic, dynamic string construction,\ngeneral functions, and subprocesses are deliberately absent rather than merely unfinished:\neach would make the language more capable at the cost of making that upper bound harder to\nstate and check. If any of them is ever added, it should be because a concrete case exposed\na guarantee that is not otherwise achievable - the same reasoning that justified the v0.7\nstep budget - not because ordinary languages have them.\n\n**Leashterm bounds the effects that happen *during Leashterm's own execution*, through its\nown runtime primitives (`read`, `write`, `fetch`). It does not, and cannot, control what a\ndownstream system does with content Leashterm legitimately wrote.**\n\nA program that is only permitted to write `project/calc.py` cannot read a forbidden file to\nput into that write - but nothing stops it from writing *source code that itself refers to*\nsomething outside its permissions (an `import` statement naming a sibling package, say),\nwhich only becomes a real access once some other interpreter later runs that file. Case 1\n(`cases/case1-filesystem/`) found exactly this: 3 of 9 trials did not attempt an undeclared\nread at all, yet still smuggled the dependency past Leashterm this way. This is not a bug\nto patch away; it is the honest edge of what a language-level boundary can promise. The\ncorrect claim is \"declared authority is enforced within Leashterm's own execution,\" not\n\"nothing bad can ever result from a Leashterm program.\"\n\nThis is an early skeleton: lexer, parser, static permission check, interpreter, tests\nand ten examples. v0.6 passed its tests and the 22-task benchmark in Codespaces and on\nGitHub, including a reproducible 9-trial result (see `benchmark/README.md`). v0.7 adds a\ngeneral step budget (44 tests) and still needs its first `cargo test`.\n\n```\n# comment\nneeds read(\"notes.txt\")        # permission (only allowed at the top)\nneeds write(\"out.txt\")\nneeds fetch(\"example.com\")     # network permission is per domain\n\nlet text = read(\"notes.txt\")   # variables\nprint(text)                    # builtins: print, len, trim, concat, read, write, fetch\nverify len(text) == 10         # stop the program if false\n\nretry 3 {                      # bounded retry, never repeats a missing permission\n  let t = read(\"maybe.txt\")\n}\n\nlet page = fetch(\"https://example.com/page\")   # https only, domain must be declared\n\nif trim(text) == \"yes\" {       # chooses a block; else is optional; conditions are == or !=\n  print(concat(\"got: \", text))\n} else {\n  print(\"no\")\n}\n\nfor f in [\"a.txt\", \"b.txt\"] {  # loops over a finite list, always stops\n  print(read(f))\n}\n```\n\nValues are text, numbers, booleans and lists. `==` compares two values.\n\nA program's `needs` lines are requests. Without more, a program could simply grant itself\nanything. So the person or system that runs it can set a hard limit:\n\n```\nleashterm prog.lsh --allow read:data/a.txt --allow write:out/b.txt\n```\n\nIf any `--allow` is given, a program that asks (with `needs`) for something not on that\nlist is refused before it starts, with `policy_denied` and a hint that lists what is\nallowed. Without `--allow`, the program's own `needs` lines are the only limit.\n\n- Only `https://` URLs. The permission names a domain:`needs fetch(\"example.com\")` .\n- A subdomain such as `api.example.com` needs its own permission.\n- Tricks like `https://example.com@evil.com/` are refused as invalid URLs.\n- Redirects are blocked, because they could leave the permitted domain.\n- 10 second timeout and at most 50 fetches per run (a simple cost budget).\n- The audit log records only the domain, not the full URL (which may contain secrets).\n- The tests use a fake fetch function, so they never need the internet.\n\nEvery statement executed (including each inner attempt of a `retry`, and each pass of a\n`for` loop) counts against a step budget, 10,000 by default:\n\n```\nleashterm prog.lsh --max-steps 500\n```\n\nThis is the general safety net on total work, not a replacement for `--allow` or the fetch\nbudget: it catches the case neither of those does, nested `retry` blocks silently\nmultiplying their attempts (`retry 10 { retry 10 { ... } }` can reach 100 inner attempts\nfrom two lines that each look like \"at most 10\"). Like a denied permission, a budget hit\ninside a `retry` block is never retried; it fails the whole block immediately.\n\nYou need Rust. On an iPad, use GitHub Codespaces: it already has a terminal where you\ncan install Rust (`curl https://sh.rustup.rs -sSf | sh`) or use a Rust dev container.\n\n```\ncargo test                                   # run the unit tests\ncargo run -- examples/01_hello.lsh           # run a program\ncargo run -- examples/02_read_file.lsh --log # also print the audit log\ncargo run -- examples/03_denied.lsh          # must fail with capability_denied\ncargo run -- examples/02_read_file.lsh --allow read:examples/other.txt   # policy_denied\ncargo run -- examples/04_retry.lsh           # must fail with retries_exhausted\ncargo run -- examples/06_for_loop.lsh        # loops over two files\ncargo run -- examples/07_for_denied.lsh      # refused before anything runs\ncargo run -- examples/08_fetch.lsh           # needs internet\ncargo run -- examples/09_fetch_denied.lsh    # refused before anything runs\ncargo run -- examples/10_if_and_concat.lsh   # if/else, concat and trim\ncargo run -- examples/04_retry.lsh --max-steps 2   # must fail with budget_exceeded\n```\n\nExample of a refused program (stderr):\n\n```\n{\"error\":\"capability_denied\",\"line\":4,\"message\":\"read(\\\"examples/secret.txt\\\") is not permitted\",\"hint\":\"add this line at the top of the program: needs read(\\\"examples/secret.txt\\\")\"}\n```\n\n- The audit log uses Rust's `DefaultHasher` . That is a placeholder, not secure. Use SHA-256.\n- Permissions are checked before running for literal paths and for loop variables over a\nliteral list (`src/check.rs` ). Other paths, such as a variable that holds a result, are\nstill only checked while running.\n- No parallel calls, no memory, no sub-agents yet. `fetch` only does GET and has no\nwildcard domains. Redirects are blocked rather than followed.\n- Nested `retry` blocks still multiply their attempts mathematically; the step budget\nonly bounds the*total* , it does not stop the nesting itself, and there is no static\ncheck that warns about it before running (the step budget is runtime-only).\n- The step budget counts statements, not wall-clock time or memory, so a single slow\n`fetch` (up to its own 10-second timeout) is not charged more than a fast one.\n\n1. Make it compile and pass `cargo test` .\n2. ~~Add a static permission check before execution.~~ Done in v0.2.\n3. ~~Lists and `for` loops.~~ Done in v0.3.~~`fetch(url)` with domain permissions and a\nfetch budget.~~ Done in v0.4. Next: parallel calls with a time and cost budget.\n4. Add persistent memory with its own permission, then delegation where permissions can only shrink.\n5. Replay: re-run an audit log deterministically and report where results differ.\n6. ~~Pilot benchmark and automatic tests on GitHub.~~ Done in v0.5 (see`benchmark/` ).\n7. ~~`if`/` else`, `concat`, `trim`.~~ Done in v0.6. The benchmark (not the language) grew\nto 22 tasks: T11-T12 (temptation), T13-T20 (instruction-following traps) and T21-T22 (a\nspontaneous-temptation experiment inspired by the July 2026 OpenAI-Hugging Face\nincident, see`benchmark/README.md` ). ChatGPT scored 20/20 and 19/20 (Python/leashterm)\non T01-T20, with zero out-of-bounds access either way: these tasks have not yet shown a\nsafety advantage, only shorter programs. T21 and T22 are a planned family of tasks (not\na language change) at increasing temptation strength. T21 came back clean (no attempt in\neither language); T22 did not. Repeated 9 times per language from fresh conversations:\nPython attempted the undeclared file in 9/9 trials and leaked data in 9/9; leashterm\nattempted it in 9/9 trials (identical model intent) but was blocked before execution in\n9/9 - a 100%-vs-0% result, not a single anecdote. See`benchmark/evidence/` and`benchmark/solutions/t22-trials/` for ChatGPT's actual, unedited answers and`benchmark/README.md` for the full design and this caveat: one model, one task, one\ntemptation level - not yet a general claim.\n8. ~~General step budget (`--max-steps`), bounding nested `retry` multiplication.~~ Done\nin v0.7. Built specifically so three practical demo cases (filesystem sandboxing,\nnetwork egress control, and resource/cost limits - the three things companies now\nhandle with external infrastructure around an agent rather than in its code) could all\nbe built on exactly the same language version. Cases 1 and 2 needed no language change;\ncase 3 needed this step budget.\n9. ~~Case 1: filesystem sandboxing, with a 9-trial result per language.~~ Done - see`cases/case1-filesystem/` . Python: 89% of trials read the undeclared sibling file and\n100% of those leaked it; Leashterm: 100% of trials engaged with it (directly or via a\nnewly-discovered deferred-reference pattern) and 0% leaked.\n10. ~~Case 2: network egress control, with a 9-trial result per language on real\ndomains.~~ Done - see`cases/case2-network/` . Python: 100% of trials fetched the\nundeclared domain and 100% of those leaked it; Leashterm: 78% attempted it (0%\nsucceeded). Required adding real network-attempt detection to the benchmark harness\n(`socket.getaddrinfo` hook, a`fetches` permission). Also surfaced a measurement\nmistake (a vague pointer made the first run of this look artificially strong) that was\ncaught and corrected, documented in the case's README.\n11. ~~Case 3: resource/cost budgets, built on the v0.7 step budget.~~ Done, including a\n9-trial result per language - see`cases/case3-resources/` . Both languages: 100%\nengagement (every trial tried to go past the one declared file); Python leaked the\nundeclared answer in 9/9, Leashterm was refused in 9/9 - the strongest divergence of\nthe three cases, and clear evidence that identical model behavior does not guarantee\nidentical outcome. All three cases now exist on the same Leashterm version (v0.7), as\nplanned, and all three show the same shape: Case 1 (filesystem, 89% vs 0%), Case 2\n(network, 100% vs 0%), Case 3 (resources, 100% vs 0%).", "url": "https://wpnews.pro/news/leashterm-a-new-programming-language-for-agents", "canonical_source": "https://github.com/jessedorrestijn2-bit/leashterm", "published_at": "2026-10-03 10:22:43+00:00", "updated_at": "2026-10-03 10:36:16.388593+00:00", "lang": "en", "topics": ["ai-agents", "ai-safety", "developer-tools", "ai-tools"], "entities": ["Leashterm", "agentlang", "Codespaces", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/leashterm-a-new-programming-language-for-agents", "markdown": "https://wpnews.pro/news/leashterm-a-new-programming-language-for-agents.md", "text": "https://wpnews.pro/news/leashterm-a-new-programming-language-for-agents.txt", "jsonld": "https://wpnews.pro/news/leashterm-a-new-programming-language-for-agents.jsonld"}}