Sentinel Vault: Runs on Atlassian AI apps A security review of Sentinel Vault, a Confluence AI review app running on Atlassian's Forge LLMs API, found the app qualifies for the "Runs on Atlassian" badge with no external data egress and no API key, but identified four areas where the implementation is weaker than its marketing claims. The review, conducted against commit e649d9f on 1 October 2026, verified Atlassian's eligibility check against the production build and ran a live AI review on a test site. The reviewer notes that while the badge confirms where data can go, Atlassian's own documentation warns such controls "do not prevent misuse of access granted to the app during installation or abuse of the app runtime. Sentinel Vault is one of the Runs on Atlassian apps that review Confluence content with AI, through the Forge LLMs API, with no API key to paste and no external data egress declared. That's the whole pitch in one sentence. Now say a request to install it has just landed in your queue, and you're the admin or the security person who has to sign off on the AI feature. Then it's also the sentence I'd tell you to distrust. At least until somebody shows you the code behind it. So this post shows you the code behind it. We read Sentinel Vault at commit e649d9f on 1 October 2026. We ran Atlassian's own eligibility check against the production build, and we read Atlassian's Forge LLMs pages the same day. We also ran one live AI review on our test site, which runs the development build. What follows is what holds, who pays for it, and how the cost ceiling actually behaves. Then come the four places where it's weaker than the marketing line. If you're deciding whether to switch the AI review on, I'd say the weak places matter more than the strong ones. Most AI features in third-party apps work the same way. You paste an API key for OpenAI or Anthropic into a settings screen. From then on, the app sends your page text to that provider. The provider's terms now apply to your content. Your security team has to review a second vendor, and somebody has to own the key. Runs on Atlassian is Atlassian's badge for apps that don't do that. Atlassian's page on the program sets out three requirements: "Apps exclusively use Atlassian-hosted compute and storage. Apps support data residency that matches data residency provided by the host Atlassian app. Customers can control external data egress for example, analytics and logs via admin controls." It also says it plainly. "Your app must not egress data, with the exception of egress for analytics purposes." You don't apply for it. In Atlassian's words, "The Runs on Atlassian badge is automatically applied to eligible apps on the Atlassian Marketplace." That matters for trust. The badge is decided by Atlassian from the deployed app, not from what the vendor writes on its listing. The Sentinel Vault listing page renders that badge today. The Marketplace's public REST API has no field for it, so we checked the listing page itself. It carries a "Runs on atlassian badge" label in its metadata. The listing shows version 6.5.0, published 30 September 2026. There's one caveat on Atlassian's own page that I'd like you to read before you treat the badge as a security guarantee: "While controls that limit external data egress are in place, these controls do not prevent misuse of access granted to the app during installation or abuse of the app runtime." The badge tells you where the data can go. It doesn't tell you the app uses its permissions well. That second question is what the rest of this post is about. What lets an AI feature live inside that badge is the Forge LLMs API. Atlassian announced it as generally available on 29 July 2026: "Today, we're announcing the Forge LLMs API is now generally available for all developers." The same announcement has the part you'll care about as an admin: "Apps that use the Forge LLMs API can qualify for Runs on Atlassian because app data stays contained within Atlassian cloud the entire time." The API reference says it again from the developer's side: "The app retains its Runs on Atlassian eligibility after the module is added." Atlassian's main page on the API puts it more simply still. "Apps using this API are badged as Runs on Atlassian." Atlassian also says it filters what goes through: "Requests to Forge LLMs undergo the same moderation checks as Atlassian first‑party AI and Rovo features. High‑risk messages per the Acceptable Use Policy are blocked." In Sentinel Vault, the declaration is four lines in manifest.yml . It names one module, sentinel-vault-llm , and one model family, claude . There's no remotes section and no external permissions block. We grepped the manifest for both and got nothing back. Beyond its Confluence scopes and a content-styles entry, it asks for read-only JSM Assets scopes used to import classification levels and app storage. None of them is an external fetch permission. The cleanest proof isn't the manifest, though. It's Atlassian's command-line check. Forge ships a forge eligibility command that tells a developer whether a deployed version qualifies. We ran it against both environments of the same app on 1 October: bash $ forge eligibility -e production The version of your app 6.5.0 that's deployed to production is eligible for the Runs on Atlassian program. $ forge eligibility -e development The version of your app 8.98.0 that's deployed to development is not eligible for the Runs on Atlassian program. - App is using a webtrigger module that can egress data Same app, two builds, two answers. The difference is the useful bit. The development build carries a web trigger we use for our own test harness. A web trigger that can return arbitrary data counts as possible egress. The production deploy script strips that module out before it deploys scripts/deploy-prod.sh runs strip-dev-modules.mjs and then runs the eligibility check . The stripped build passes. You can see the same care in a newer feature. Since mid-September the app has a configuration REST API, and it was built deliberately as a static web trigger. The comment in the manifest says why: "a static trigger cannot egress anything the manifest did not spell out, which keeps 'Runs on Atlassian'." The static trigger was in the design from the first commit. That's the kind of decision I like to see a vendor make up front. The implementation commit also closed an internal review's findings, one being that the roster, AI settings and rule text are "never mirrored" out through it. Because the model runs on Atlassian's side, there's nothing for you to configure in terms of credentials. The AI section of Sentinel Vault's default configuration has a switch, a model, thresholds and a budget. It has no key field. The comment at the top of the app's LLM client explains the architecture in two lines. The admin screen gives you the consequence in one: "Token usage is billed to this app's Forge account, so AI is off by default and limited to Claude Haiku." That last sentence is the interesting one, so let me take it apart. Marketing pages say "off by default" easily. We wanted it from the code. The default configuration in logic.js sets the validation engine's master switch to enabled: false . AI has its own switch, ai: { enabled: false, ... } , which is independent of the master switch and also off. Then every place that could start an AI review checks ai.enabled === true before doing anything. That's the manual review request, the workflow gate, and both entry points in the background worker. A fresh install doesn't call the model at all until an admin turns it on. The manual review is also admin-only. The code works out the space from the page itself instead of trusting what the caller says. Anyone else is refused with "Only an admin of this page's space can run an AI review". So a reader of a page can't spend the budget by clicking a button. An editor can, though. In most setups, requesting a move into an AI-gated state queues a review, and so does re-requesting it after an edit. The app sends text only. Anything multimodal is flattened to text before the call. The page text is capped at 40,000 characters by default, and the answer at 4,096 output tokens. Atlassian's own limits per installation are a 200,000-token context, 100 requests a minute and 50,000 tokens a minute per model. A single page review sits comfortably inside them. The call itself runs in a background queue, not in the request you make. The worker's comment explains why: "The Forge LLM call can exceed the 25s resolver limit, so it runs on the ai-validation-queue 120s function ." You ask for a review, and the verdict lands a little later. Here's the fact that shapes everything else. Atlassian's pricing page for Forge LLMs is explicit: "No free usage allowance: The Forge LLMs API does not include a free monthly usage quota. All token usage is billed. Forge LLMs usage is charged to the developer of the Forge app and counted toward your Forge monthly bill." The developer of the app is us. Every token your site spends on a Sentinel Vault review is a line on LeanZero's Forge bill, not on your Atlassian invoice. The first line of the app's LLM client says the same thing in code: "Token costs bill to the app vendor's Forge bill, so we enforce a Haiku-only policy see isForgeLlmModelAllowed at every layer." Atlassian's pricing page puts Haiku 4.5 at 10 credits per million tokens. That works out at $1 per million input tokens and $5 per million output tokens. Sonnet 4.5 is $3 and $15. Opus 4.6 is $5 and $25. We don't have a measured bill for Sentinel Vault reviews to show you. The app doesn't surface usage yet, and we didn't read the counters on our installs. So we won't quote a cost per review. The rate card is the honest number. The way I see it, this changes the incentive in a useful way for you. The vendor has every reason to keep the model small and the calls few, because the vendor is paying. A bring-your-own-key app has the opposite incentive: your key, your bill. If you'd like the other side of that trade laid out for a Jira app, our Forge LLM tutorial for Jira https://leanzero.net/tutorials/build-an-llm-powered-atlassian-forge-app-for-jira?utm source=devto&utm medium=referral&utm campaign=crosspost builds AI validators on the same API. The manifest declares the model family claude . Atlassian's models page lists three tiers under that family: "Forge LLMs supports Claude models across three tiers: Haiku, Sonnet, and Opus." So the platform would let the app call Opus. The narrowing to Haiku is our policy, not Atlassian's. It's enforced in code at four separate points. The rule itself is a one-line regular expression, /haiku/i , on the model id. It's applied when the admin screen lists the models it offers, so only Haiku appears. It's applied when a configuration is saved, so a non-Haiku id is never stored. That includes saves through the new configuration REST API. It's applied in the background worker, in both of its entry points one marked with the comment "cost backstop" . And it's applied in the chat adapter itself, which logs that a model "not allowed" was clamped. One detail about direction, because it matters if you're scripting configuration. A non-Haiku model id isn't rejected with an error. It's quietly replaced with claude-haiku-4-5-20251001 . If you set a Sonnet id through the API, the save succeeds and the app runs Haiku anyway. Our own design note on the cost ceiling describes this as three layers. The code at today's commit has four. The worker check was already there when the note was written, and the note missed it. I'm only mentioning it so that if you compare the two, you know which one to believe. The client retries transient failures. That means HTTP 429, 408, any 5xx, or a matching error message. It makes up to four attempts in total, waiting 400, 800 and 1,600 milliseconds between them. There's a Math.min 2000, ... cap in that delay calculation that can never actually bind, because the largest delay it computes is 1,600. It's dead code. Harmless, and a fair sign that nobody has had to tune it. The Forge LLMs chat call has no structured-output option. So the app can't ask the model for JSON the way some APIs allow. Instead it adds an instruction to the system message: "Respond with ONLY a valid JSON object. No markdown fences, no surrounding prose, no explanation outside the JSON." Then a tolerant parser recovers the answer. It strips code fences and tries a plain parse. Next it falls back to the outermost object or array, repairs unescaped quotes, and repairs a truncated answer. It returns nothing instead of throwing. Its test file passes 11 of 11 on today's code. What happens when that still fails is the part I like. In a workflow gate, if the model's answer was cut off by the token limit, the code detects it and never grants a pass. The verdict is "AI review was cut off too long — please retry." A manual review doesn't check this. It keeps whatever the parser salvaged from the cut-off answer. If the answer can't be parsed at all, the app records it in the audit trail and posts no comment. The comment in the code reads "fail-closed — never fabricate". Atlassian's API reports input and output tokens on every response, and the worker adds them to a counter after each call. Sentinel Vault uses the AI in two ways. A space admin can ask for a review of a page. And since July, a workflow state can require an AI review before a page is allowed to move into it. That second one means correcting something we say ourselves. Our Sentinel Vault product page https://leanzero.net/portfolio/sentinel-vault?utm source=devto&utm medium=referral&utm campaign=crosspost says "The deterministic engines do the enforcing — the AI advises." That was true when it was written. Since the July release that added transition conditions, a workflow state can carry an entry condition, requireAi . When it's on, the AI's verdict can block a page from entering that state. It's off by default, and you choose the threshold low, medium or high, default medium . But once you switch it on, the AI is enforcing, not advising. The code is the thing to trust here. The gate reviews the pinned version of the page, the one being approved, not whatever the latest edit is. If the model call fails, the page can't be read, or the answer can't be parsed, the gate ends in a terminal failure. Never a pass. Since late September, space admins who act as approvers also count towards the quorum on an AI-gated approval. If you haven't seen the rest of the app, our post on why you can't block a Confluence save and what Sentinel Vault does instead https://leanzero.net/blog/sentinel-vault-detect-and-restore-confluence?utm source=devto&utm medium=referral&utm campaign=crosspost covers the detect-and-restore engine that does the non-AI enforcing. Our product page says the AI is "capped by a monthly token budget you set". That's true. It's also the sentence we'd push back on hardest if we were reviewing this app for our own site. All four of the following are open in the code today, and our own design note lists them. The default budget is zero, and zero means unlimited. The configuration line reads monthlyTokenBudget: 0, // 0 = unlimited , and the admin screen starts at the same value. If you switch AI on and don't set a number, there's no ceiling. The budget is counted per space. The counter key is built from the space key and the month, ai-usage-