Quick answer: an AI context layer starts going stale the day you write it. Code changes, business rules change, and the correction someone made in a Slack thread scrolls away. Four feedback loops keep it current: local review, CI, scheduled audits, and channel feedback. Each one catches a different kind of drift, and they all end the same way, with an owner, a tracked request, and a reviewed change. At no point does the agent silently edit context.
This comes from the September Build Lab session on context for agents. It is phase four of onboarding an agent like a new hire, where the agent is hired, real teams use it, and maintenance becomes part of the job.
Why context goes stale #
Nobody has to be careless for this to happen. Three ordinary things wear it down.
The code moves. A backend release renames an order status, adds a column, or changes a unit, and the asset description, the check, and the glossary all still describe last month.
The business moves too. Finance decides refunds come out of net revenue in the period they happen, and the semantic model still says otherwise.
And corrections evaporate. Someone tells the agent "that spike is expected on the first of the month". That sentence lives in a thread, so the next person and the next agent make the same mistake.
Stale context is worse than none, because it looks correct. Agent memory does not fix this either. Memory is private to one session or one person, nobody reviewed it, and a new session or a different tool loses it. Write it where the team reads it, which is the repo.
The rule for every loop #
Every loop has to produce the same three things:
- An owner - the person who can say what the definition should be.
- A tracked request - an issue or a pull request. A chat message does not count.
- A reviewed change - merged by a person, in the same surface you use for code.
This is the usual permission rule for agent tools, made concrete. The agent gets broad read access so it can see drift anywhere, and narrow write access so the most it can do is propose.
Loop 1: local review #
The cheapest loop runs while someone is working. When a person corrects the agent during a task, the agent updates the context in that same task, in the same pull request as the fix.
Write this into AGENTS.md as an instruction:
## When you are corrected
- Find the context that misled you: the asset description, a column, a check,
the glossary, or the semantic model.
- Update it in the same pull request as the fix, and say what you changed and why.
- Ask why before writing it down. "Stop flagging it" and "this is expected on the
first of the month" are different rules.
- If the correction changes a metric definition, do not write it. Flag the owner.
This catches errors before they reach anyone else. A correction today is context every agent has tomorrow.
Loop 2: CI on every pull request #
The second loop reviews changes before they get a chance to break something. On every pull request, you validate the context the same way you validate the code.
name: context
on: pull_request
jobs:
context:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: bruin-data/setup-bruin@main
- name: Write Bruin config
env:
BRUIN_CONFIG: ${{ secrets.BRUIN_CONFIG }}
run: printf '%s' "$BRUIN_CONFIG" > .bruin.yml
- run: bruin validate .
- run: bruin format --fail-if-changed .
- run: ./scripts/context_audit.sh
bruin validate and bruin format --fail-if-changed are Bruin CLI commands. The last step is not. It is a script you write, or an agent you call, that compares what the asset says with what the data shows. For example, it can flag a status value that appears in the warehouse but is missing from the column's accepted_values check.
When it finds something, it drafts a candidate change and explains it in plain language, so the reviewer does not have to reverse-engineer a diff:
Illustrative mockup. The custom-context-audit check is custom logic, not a built-in Bruin command.
I call it a candidate on purpose. A value observed in the warehouse is not automatically a valid business value. refunded showing up in stg_orders might be a real new status, or it might be a bug in the app. The owner confirms the definition before merge.
Loop 3: scheduled audits #
CI only sees what changes in the repo. Drift that starts somewhere else (in the source app, in the warehouse, in who owns what) needs something that goes looking for it on a schedule.
In Bruin Cloud this is a scheduled agent: an existing AI agent given a task that runs on a recurring schedule. Each run uses the agent's connections, integrations, and CLI access, and posts to Slack, Teams, WhatsApp, or the Bruin Cloud chat. The audit is the task you give it, meaning what to scan, what counts as drift, and how to report it.
A useful audit sorts what it finds into two buckets:
Illustrative mockup. The audit logic and the message format are custom; Bruin does not decide which context changes to make.
- Suggested changes. Mechanical drift, like a new value, an owner change, or a missing column, drafted as pull requests through your CI or GitHub automation.
- Needs a human. Anything that changes business meaning, such as a revenue definition that may now exclude refunds. The agent flags it and stops there.
That separation is the trust model. The scheduled agent delivers the report, and the reviewable change goes through the same path as any other code.
For pure schema drift there is a simpler version of this loop. A weekly job that runs bruin import database and bruin ai enhance and opens a pull request with the diff turns new tables and columns into a reviewable change. ai enhance is additive, so it fills gaps without overwriting what people wrote. The AI context layer post covers that setup.
One cost to watch: a scheduled audit runs queries whether or not anything changed, and on most bills the warehouse spend matters more than the model tokens. Scope the audit to the assets that matter, and pick a cadence that matches how fast they change.
Loop 4: channel feedback #
The last loop catches what no audit can, which is knowledge that only comes up when a person reacts to an answer.
Illustrative mockup. Issue #482 is fictitious.
The order of these steps is what keeps it safe:
- Report. A data consumer says
revenue_dailylooks off and may still count refunded orders. - Verify. The agent checks the SQL, the canonical definition of net revenue, and whether a known-question eval fails. At this stage the report is a lead, nothing more.
- Draft an issue. Configured automation opens an issue with the evidence and a proposed fix, written as a question for the owner rather than a confirmed bug.
- Review. The metric owner decides what the definition should be.
- Pull request. The fix lands with the glossary entry or semantic model updated, plus a check so the same gap cannot quietly open again.
The issue should say what was checked and what is being proposed. If it presents an unverified cause as fact, that is exactly how a context layer collects confident mistakes.
What each loop catches #
| Loop | Triggered by | Catches |
|---|---|---|
| Local review | a person correcting the agent mid-task | errors before anyone else sees them |
| CI | every pull request | invalid context, formatting drift, definitions that disagree with the warehouse |
| Scheduled audit | a cadence you choose | drift from outside the repo: new values, new tables, ownership changes |
| Channel feedback | a user reacting to an answer | wrong business meaning nobody wrote down |
CI and scheduled audits keep the context matching reality. Channel feedback keeps it matching what the business means. You want all four, because they catch different failures.
Recycle while you are there #
Keeping context current also means taking things out. Every loop is a chance to delete something.
If a loop finds two sources that disagree, pick one home for the fact and remove or rewrite the other copy (a conflict is worse than a gap, and The Best Context Is No Context explains why). If an audit turns up tables, reports, or dashboards nobody uses, remove them rather than documenting them. And if a page is wrong, delete it. A wrong page is worse than no page.
Measure that it is working #
Keep the known-question set from the internship phase in the repo and rerun it after context changes. If accuracy holds or goes up, the loops are doing their job. If it drops after a merge, the last change made the context worse, and you know exactly which pull request to look at.
Every incident and every correction should leave something behind: a check, a description, or a rule. When one leaves nothing, expect to see it again.
FAQ #
How do I keep AI agent context up to date?
Run four feedback loops. Local review: when an agent is corrected during a task, it updates the context in the same pull request. CI: every pull request validates the context and flags drift between the asset definitions and the warehouse. Scheduled audits: an agent checks for drift on a cadence and reports what changed. Channel feedback: when a user reports a wrong answer in Slack or Teams, it is verified and turned into a tracked issue. All four end in a reviewed change with an owner.
Should an AI agent edit its own context automatically?
No. The agent can detect drift and draft a change, but it should never silently edit context. It proposes a pull request or an issue in the same place you review code, and a person who owns the definition approves it. A value observed in the warehouse is not automatically a valid business value.
Why not just use the AI agent's memory for context?
Agent memory is private to one session or one person, unversioned, and easy to lose with a new session or a different tool. Nobody reviewed it, and you cannot tell when it was learned or whether it is still true. Context written into the repository is reviewed, versioned, and read by every engineer and every agent.
Can Bruin run a scheduled context audit?
Bruin Cloud has scheduled agents: an existing AI agent given a task that runs on a recurring schedule, using that agent's connections, integrations, and CLI access, and posting results to Slack, Teams, WhatsApp, or the Bruin Cloud chat. The audit logic itself is yours to write as the agent's instructions; Bruin does not decide which context changes to make.
What should a context audit check?
Mechanical drift first: tables or columns in the warehouse with no asset definition, values present in the data but missing from accepted_values checks, owners who left, and descriptions that contradict the query. Then flag, but do not fix, anything that changes business meaning, such as a metric that may now include refunds. Those go to the metric owner.
Related Reading #
- How to Onboard an AI Data Agent Like a New Hire - the four phases that come before these loops.
- The Best Context Is No Context - what to write down, what to leave out, and why one fact needs one home.
- How to Run Data Pipelines in CI/CD - the CI setup this loop builds on.
- What Is a Self-Healing Data Pipeline? - the same write-it-back loop, applied to incidents.