{"slug": "the-best-context-is-no-context", "title": "The Best Context Is No Context", "summary": "Writing extensive context documentation for AI agents is counterproductive because most of it duplicates facts the system already exposes through configs, logs, metadata, and queries, according to a September Build Lab session on agent context. The session recommends a three-part approach — reuse inherent context from tools, reduce explicit context to only missing meaning such as business definitions and ownership, and recycle by resolving conflicts so each fact has one source of truth. The guidance applies across stacks including Bruin, dbt, Airflow, and BigQuery.", "body_md": "**The short version:** most context written for AI agents is a copy of something the system already says. The copies drift, the agent finds two versions, and it picks one. So treat context like any other code and keep it small. Reuse what your code, configs, and logs already show, reduce what you add to the meaning they cannot show, and recycle anything that has gone stale or contradicts its source. The title sounds backwards, but the idea is simple: do not document what the system already states clearly.\n\nThis comes from the September [Build Lab](https://getbruin.com/build-lab/) session on context for agents. It is also the rule that decides what you write in every phase of [onboarding an agent like a new hire](https://getbruin.com/blog/onboard-ai-data-agent-like-a-new-hire/).\n\n## [The instinct to write everything down](#the-instinct-to-write-everything-down)\n\nWhen an agent gets something wrong, the reflex is to add more context. Another paragraph in the README, a longer `AGENTS.md`, maybe a wiki page that explains the pipeline from scratch.\n\nSome of that helps. Most of it restates facts the agent could already read, and every restated fact is a second copy that has to stay in sync with the first. Nobody keeps it in sync. Six months later the README says the pipeline runs every six hours, `pipeline.yml` says hourly, and the agent has to guess which one you meant.\n\nThe agent usually has plenty of text. What it is missing is meaning, and a clear answer to \"which of these is true?\"\n\n## [Reuse: start with the context you already have](#reuse-start-with-the-context-you-already-have)\n\nA lot of context is inherent. It exists because the system runs, and since it is what actually executes, it is the most accurate description you have.\n\nConfigs describe schedules, retries, dependencies, and where alerts go. Logs show what ran, what failed, and what the error said. Metadata exposes schemas, types, owners, tags, and lineage. The data itself shows which values actually appear and how tables relate. And the query shows how a table is built: which columns feed it, which joins and filters shape it, and where the business logic lives.\n\nThat holds whatever the stack is - Bruin, dbt, Airflow, BigQuery, or any mix of them. So before you write a sentence of context, ask whether a tool can already answer the question. If it can, point the agent at the tool, give it access, and keep the explanation out of the context layer.\n\n## [Reduce: add only the missing meaning](#reduce-add-only-the-missing-meaning)\n\nPicture two circles. One is inherent context, the facts the system already exposes. The other is explicit context, which is everything a person wrote down. Where they overlap is waste: prose that repeats a fact the tools already show. Cut that down to the minimum viable information.\n\nWhat does earn a line of explicit context:\n\n- Business definitions. \"AOV\" is a familiar acronym, but the agent still needs your definition of it: which orders count, and whether refunds, taxes, and shipping are excluded.\n- Decisions, like why de-duplication happens once in staging, so nobody adds a second pass downstream.\n- Ownership. A real person to escalate to means the agent asks instead of guessing.\n- Exceptions, such as \"the spike on the first of the month is expected.\"\n- The why behind things. Anything a new engineer asks in their first week that the SQL cannot answer.\n\nA gap prompts a question; an unclear definition produces a confident mistake. That difference is the reason explicit context is worth writing at all, and the reason it should be short.\n\n## [Recycle: resolve conflicts and overlap](#recycle-resolve-conflicts-and-overlap)\n\nRecycle is the maintenance loop, and it runs again every time the code or the business changes:\n\n1. **Identify** where two sources overlap or disagree.\n2. **Choose** one source of truth for that fact.\n3. **Remove** the other copy, or rewrite it so it points at the source instead of competing with it.\n\nThis is not a one-time cleanup. Code changes, business rules change, and documentation goes stale without anyone noticing. Deleting a wrong page counts as a context improvement, and it is often a bigger one than writing a new page.\n\n## [Why a conflict is worse than a gap](#why-a-conflict-is-worse-than-a-gap)\n\nWhen context is missing, the agent hits a visible unknown. A well-instructed agent says so (\"I could not find a definition of active customer\"), you fill the gap, and you move on.\n\nWhen two sources disagree, the agent has two plausible answers and nothing telling it which to trust. So it picks one and sounds sure about it, and the wrong answer looks justified because it came from your own documentation. Nobody notices until a number is wrong in a meeting.\n\nThree rules keep this under control:\n\n- **One fact, one home.** Each fact has exactly one authoritative place, and everything else references it.\n- **State the precedence.** Where scopes overlap, narrower scope wins, and you write that rule down where the agent reads it. In[Bruin](https://github.com/bruin-data/bruin) , a column can`extends` a glossary term to inherit its type and description, and anything set explicitly on the asset overrides the glossary default.\n- **Keep the open question visible.** If a conflict cannot be resolved today, record it with an owner. Do not let the agent quietly pick a side.\n\nThis is one way to split the facts so each has a single home:\n\n| Fact | Its one home | \n|---|---|\n| Schedule, retries, alert channels | pipeline config ( `pipeline.yml` ) | \n| Why the pipeline exists, who consumes it | pipeline `README.md` | \n| What one row represents, owner, tier | the asset definition | \n| How the table is built | the query | \n| Valid values, keys, invariants | column checks and custom checks | \n| Shared terms like \"order id\" | the glossary | \n| Metric definitions like revenue or AOV | the semantic layer | \n| Decisions, runbooks, project plans | documentation, linked from the repo | \n| Where each of the above lives, and what wins | `AGENTS.md` and its context map | \n\n## [Shared meaning belongs in one layer](#shared-meaning-belongs-in-one-layer)\n\nThe facts behind most conflicts are business meanings: what \"active customer\", \"net revenue\", or \"order id\" mean. They show up in every pipeline, so they get redefined in every pipeline.\n\nGive them one layer. In Bruin that is two files at the repository level. A [glossary](https://getbruin.com/docs/bruin/getting-started/glossary.html) defines shared domains, entities, and attributes, so columns extend a definition instead of repeating it (the glossary is marked beta in the docs). A [semantic layer](https://getbruin.com/docs/bruin/core-concepts/semantic-layer.html) defines reusable metrics and dimensions in YAML under `semantic/`:\n\n```\n# semantic/orders.yml\nschema: v1\nname: orders\nlabel: Orders\ndescription: One-time storefront order metrics.\n\nsource:\n  table: shop_report.orders_enriched\n\ndimensions:\n  - name: order_date\n    type: time\n    granularities:\n      month: date_trunc('month', order_date)\n  - name: country\n    type: string\n\nmetrics:\n  - name: revenue_usd\n    expression: sum(amount_usd)\n  - name: order_count\n    expression: count(distinct order_id)\n  - name: average_order_value_usd\n    expression: \"{revenue_usd} / {order_count}\"\n```\n\nIf you want agents to build dashboards and charts, add a semantic layer. The agent then reuses `revenue_usd` instead of rebuilding revenue from raw columns on every question, slightly differently each time.\n\n## [Fix the model before adding context](#fix-the-model-before-adding-context)\n\nContext can explain the gaps around a sound system. It cannot repair a system nobody understands. If a senior data person cannot navigate your models, read the code, and understand the business, an agent cannot either. It will still write fluent SQL, but fluency is not understanding.\n\nSo before writing more context, check the model itself:\n\n- Add a semantic layer, so metrics, dimensions, entities, and relationships each have one reusable meaning.\n- Model with purpose. That means a stated grain, tiers that fit the team, readable SQL, and no logic duplicated across layers.\n- Get rid of stale data. Audit and delete unused tables, reports, and dashboards rather than making the agent search through them.\n- Delete outdated pages. A wrong page is worse than no page.\n\n## [The evidence for less](#the-evidence-for-less)\n\nThe best published number I know of points the same way. Anthropic's data team [reported](https://claude.com/blog/how-anthropic-enables-self-service-data-analytics-with-claude) that without skills, Claude's accuracy on their analytics evals \"didn't exceed 21%\", and adding skills got it \"consistently above 95% in aggregate\". Giving the agent grep access to thousands of historical dashboard, transformation, and notebook SQL files moved accuracy \"by less than a point in either direction\".\n\nIt is one internal evaluation of a broader system, and I would not treat it as a benchmark for your stack. Still, the signal is clear enough. More raw SQL barely moved accuracy, and the curated, written-down knowledge did the work. So I would aim for the least context that closes the real gaps, with every fact in one place.\n\n## [FAQ](#faq)\n\n### [What does \"the best context is no context\" mean?](#what-does-the-best-context-is-no-context-mean)\n\nDo not write down what the system already states clearly. A table schema, a pipeline schedule, a query log, or the SQL itself is already context an agent can read. Extra prose only helps when it adds meaning the system cannot show on its own: business definitions, decisions, ownership, exceptions, and why a rule exists.\n\n### [Why is conflicting context worse than missing context for an AI agent?](#why-is-conflicting-context-worse-than-missing-context-for-an-ai-agent)\n\nA missing definition leaves a visible gap the agent can report or ask about. Two conflicting definitions give it two plausible answers, so it picks one and sounds confident, and the wrong answer looks justified because it came from your own documentation. Give each fact one authoritative home and state a precedence rule where scopes overlap.\n\n### [What context should I write for an AI data agent?](#what-context-should-i-write-for-an-ai-data-agent)\n\nOnly what closes a real gap: what a metric like AOV means in your business, which orders count as revenue, who owns a table, why a pipeline is shaped the way it is, and the exceptions people know but never wrote down. Leave schedules, column types, row counts, and dependencies to the configs and schemas that already hold them.\n\n### [Where should each piece of context live?](#where-should-each-piece-of-context-live)\n\nOne fact, one home. Run behaviour in the pipeline config, the reason a pipeline exists in its README, table grain and ownership in the asset definition, shared terms in a glossary, metric definitions in a semantic layer, valid values in checks, and decisions or runbooks in documentation that the repo links to. Everything else should point at that home rather than restate it.\n\n### [Do I need a semantic layer for AI agents?](#do-i-need-a-semantic-layer-for-ai-agents)\n\nIf you want agents to build dashboards and charts or answer metric questions, yes. A semantic layer gives the agent a reusable, reviewed definition of each metric and dimension, so it reuses revenue instead of reconstructing it from raw columns on every question, slightly differently each time.\n\n## [Related Reading](#related-reading)\n\n- [How to Onboard an AI Data Agent Like a New Hire](https://getbruin.com/blog/onboard-ai-data-agent-like-a-new-hire/) - the four-phase plan this rule is applied in.\n- [How to Keep an AI Context Layer From Going Stale](https://getbruin.com/blog/keep-ai-context-layer-current/) - the recycle step, automated with CI, scheduled audits, and channel feedback.\n- [How to Build an AI Context Layer for Your Data Warehouse](https://getbruin.com/blog/build-ai-context-layer-data-warehouse/) - generating the inherent context with two commands.\n- [What Is Dashboards as Code?](https://getbruin.com/blog/what-is-dashboards-as-code/) - why a semantic layer matters once agents build the charts.", "url": "https://wpnews.pro/news/the-best-context-is-no-context", "canonical_source": "https://getbruin.com/blog/best-context-is-no-context/", "published_at": "2026-10-01 00:00:00+00:00", "updated_at": "2026-10-05 16:17:00.034736+00:00", "lang": "en", "topics": ["ai-agents", "artificial-intelligence", "ai-tools"], "entities": ["Build Lab", "Bruin", "dbt", "Airflow", "BigQuery"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/the-best-context-is-no-context", "markdown": "https://wpnews.pro/news/the-best-context-is-no-context.md", "text": "https://wpnews.pro/news/the-best-context-is-no-context.txt", "jsonld": "https://wpnews.pro/news/the-best-context-is-no-context.jsonld"}}