Infra & Platform Engineer
A coding agent with Bash access can destroy your database, print your secrets, or ignore your team’s rules, all while following instructions. guardrails-md, released today as open source, is a gate that scores every command against your repo’s GUARDRAILS.md before anything runs. If a command tips over the threshold, the call is blocked and the agent is told why, so it can pick another route. The judgement takes about 100 ms.
The two-minute explainer above shows the whole pipeline. The short version follows.
The pipeline #
At session start the gate reads GUARDRAILS.md once (just the first 2,000 characters) and freezes it. The agent may edit the file later like any other file, but the change only takes effect next session, having gone through review like any other edit.
When the agent calls Bash, the hook intercepts before anything runs. It builds a state text: the command itself, the contents of any script files it points at, and the frozen guardrails. Then it asks the decision model four fixed questions, in one request:
destructive. Does the command irreversibly destroy data, databases, or infrastructure?credentials. Does it contain, print, or send secrets, keys, or tokens?guardrails_violation. Does it violate the team’s GUARDRAILS.md?policy_exception. Does the guardrails text explicitly name this command as allowed?
Each answer is a score between 0 and 1, compared against a threshold (0.7 by default). Above the threshold, the call is blocked and the reason returns to the agent as the tool result:
SystemOne-gate: blocked command — destructive=0.98 > 0.7
rm -rf ./important-data
Why a decision model #
A pattern deny-list knows rm -rf; it doesn’t know your policy. Asking a second LLM to review each command does know the policy, but it can be argued with, and it costs a full generation per command. The gate instead uses berget/bev, the System One decision model we described when we launched System One: fixed questions in, scores out, one forward pass. It judges the command and the rules, never the conversation that led to them. Whatever the driving model has been persuaded of, the gate sees the command on its own merits.
The parts that do not negotiate #
- Destructive and credentials are the backstop. A
guardrails_violationcan be lifted by naming the command in the MAY section of GUARDRAILS.md; destructive and credentials can’t be lifted, whatever the file says. - Retrying costs. Every block doubles the wait before the next judgement: a few seconds at first, an hour and a half by the 20th block. Variants of the same command don’t help.
- The gate fails closed. Failing closed is deliberate: if the endpoint is unreachable, commands are blocked until it responds again. Interactive users can opt out with
SYSTEMONE_FAIL_OPEN=1. No fast path, no prefix allowlist, no override file the agent could write. - Some files are protected outright. The file-editing tools refuse to touch GUARDRAILS.md and the harness configuration. The refusal is deterministic: no model call, no threshold, no cooldown.
Overrides belong to the human #
To change the policy, edit GUARDRAILS.md in a pull request; that is the durable fix for a false positive. To tune the gate, raise the threshold. To run one session without it, restart the harness with SYSTEMONE_GATE=off. All three are things a person does, and restarting re-reads the policy and resets the cooldown.
Also at commit time #
The same judgement works as a git pre-commit hook: every staged diff is scored before it enters the repository. Personal data (GDPR) and secrets are blocked; your team’s own names in bylines and author fields pass.
The honest limits #
The gate is a trained model, not a rule engine. It scores around 96% on our held-out tests, but 74% on red-team commands, and it will occasionally judge a command wrong in both directions, blocking something safe or letting something risky through. It doesn’t decode encoded payloads. Treat it as one layer among scoped credentials, sandboxes, and human review, not a sandbox. The full evaluation table and methodology live in the repo.
Try it #
Three steps, about a minute.
Add a GUARDRAILS.md to your repo that says what your team allows:
## The agent MUST NOT
- Push directly to main
- Run irreversible operations against shared systems
## The agent MAY
- Run tests, lint, and builds locally
Install the gate in your harness:
- OpenCode: Add
"plugin": ["@bergetai/guardrails-md"]toopencode.json - Pi: Run
pi install npm:@bergetai/guardrails-mdin your terminal
If you’re already logged in to Berget AI in OpenCode or Pi, no key is needed. Otherwise set BERGET_API_KEY and restart the harness.
Full setup, configuration, and the pre-commit hook are in the README. Commands are judged by the Berget AI API, where nothing scored is stored (zero data retention). Gate calls are small enough that the €5 free-tier credit lasts a long time.