{"slug": "schema-guard-stop-ai-agents-from-inventing-column-names-in-sql", "title": "Schema-guard, stop AI agents from inventing column names in SQL", "summary": "Schema-guard, an open-source tool that snapshots real table and column names into a repo file (.schema-guard/schema.json) and validates AI coding agents' SQL against it before the SQL runs, produced 48 of 48 runnable files across 4 requests × 3 arms × 3 runs on Claude Code, versus 0 of 24 without the snapshot, according to the project's published eval. In the eval, Claude Haiku 4.5 and Sonnet 5 both failed all 12 baseline files on missing column names, while the snapshot-plus-hook arm ran 12 of 12 files for each model, with 12 of 12 correct for Haiku and 11 of 12 for Sonnet. The hook arm cost about the same as baseline ($0.65 vs $0.58 for 12 Haiku runs; $1.61 vs $1.61 for Sonnet), and the author discloses the eval is small and synthetic with 3 runs per cell and two grader references added after reading runs.", "body_md": "**Your AI agent stops inventing column names.**\n\nCoding agents write SQL against the schema they *think* you have, based on your README, an old query, or a naming\nconvention. Then it fails in CI, in a dashboard, or at 2am. Snowflake's own developer blog ran a whole post on this\nin September 2026 ([My coding agent won't stop hallucinating table columns](https://www.snowflake.com/en/developers/blog/coding-agent-hallucinating-table-columns/)).\n\nschema-guard keeps a snapshot of your real tables and columns in the repo (names and types only, no data, no\ncredentials) and checks the agent's SQL against it **before it runs or lands in a file**:\n\n```\nschema-guard: this SQL names things that are not in the schema snapshot (.schema-guard/schema.json, taken 2026-09-29T01:35:07Z):\n- `analytics.customers` has no column `country`. Did you mean: `country_iso2`?\n- `analytics.orders` has no column `customer_id`. Did you mean: `cust_id`, `order_id`?\n- `analytics.customers` has no column `id`. Did you mean: `cust_id`?\nFix the names and try again. ...\n```\n\nThat's a real denial from the eval below: Claude Haiku 4.5 writing `models/revenue_by_country.sql` from a README\nthat describes last year's schema. The agent reads the denial, fixes its SQL and moves on. You never see the broken version.\n\n**Setup.** An analytics repo whose README describes an older schema (`customer_id`, `created_at`, `country`),\nplus two up-to-date models that use a few of the real names. The agent can read and write files but can't reach\nthe warehouse. That's the situation in Snowflake's post: the agent has the repo, not the account. Each request\nasks for a new SQL model, and afterwards the grader runs every file the agent wrote against the real DuckDB\nwarehouse. 4 requests × 3 arms × 3 runs, on Claude Code.\n\n| Arm | Haiku 4.5: fails on a missing name / runs / correct | Sonnet 5: fails / runs / correct | \n|---|---|---|\n| baseline (repo only) | **12** / 0 / 0 of 12 | **12** / 0 / 0 of 12 | \n| rule (snapshot + one line in CLAUDE.md) | 0 / 12 / 11 of 12 | 0 / 12 / 12 of 12 | \n| hook (snapshot + hook, no instruction) | 0 / **12** /**12** of 12 | 0 / **12** / 11 of 12 | \n\nWhat that means:\n\n- **Without a snapshot, neither model wrote one working file (0 of 24).** Both trusted the README, and even when they\ncopied real names from the existing models they mixed them with stale ones.\n- **With the snapshot, every file ran (48 of 48).** The 2 wrong answers are logic errors, not names: Haiku started\nweeks on Sunday, and Sonnet counted the last days of 2025 in the first week.\n- **The hook is the safety net for agents that don't go looking.** Haiku was denied in 10 of its 12 hook runs and\nfixed the names on the first retry every time. Sonnet found`.schema-guard/schema.json` by itself and was never\ndenied. A one-line rule gets the same result if the agent follows it; the hook doesn't depend on that, and it also\ncovers ad-hoc queries and MCP tools.\n- Cost: the hook arm cost about the same as baseline ($0.65 vs $0.58 for 12 Haiku runs; $1.61 vs $1.61 for Sonnet).\n\n**Be skeptical of this:** the world is small and synthetic, and the stale README is designed in (docs drift is\nnormal, but I chose how far). There are 3 runs per cell. Two grader references were added after I read runs:\n\"net revenue\" net of refunds (it changed 2 Haiku grades, one rule run and one hook run), and listing all 52 weeks\nwith zeros (5 Sonnet grades). Both are disclosed in [evals/scenarios.py](https://github.com/idk-arsh/schema-guard/blob/master/evals/scenarios.py), and every run's SQL is in\n[evals/results/](https://github.com/idk-arsh/schema-guard/blob/master/evals/results). Rerun it: `cd evals && python run_eval.py --model <model> --runs 3`.\n\n```\npip install \"schema-guard[duckdb] @ git+https://github.com/idk-arsh/schema-guard\"\nschema-guard snapshot --dbt target        # or --duckdb, --bigquery, --snowflake, --databricks, --url, --ddl, --csv\ngit add .schema-guard/schema.json\n```\n\nThen use it however your team works. All of these read the same snapshot:\n\n| Where | How | \n|---|---|\n| **Claude Code** (hook) | `/plugin marketplace add idk-arsh/schema-guard` then`/plugin install schema-guard` , or`schema-guard install` to add it to`.claude/settings.json` | \n| **Cursor, Claude Desktop, VS Code, Windsurf** (MCP) | `{\"command\": \"uvx\", \"args\": [\"--from\", \"git+https://github.com/idk-arsh/schema-guard\", \"schema-guard-mcp\"]}` . Tools:`list_tables` ,`describe_table` ,`search_columns` ,`check_sql` | \n| **pre-commit** | `- repo: https://github.com/idk-arsh/schema-guard` /`rev: v0.1.0` /`hooks: [{id: schema-guard}]` | \n| **CI** | `schema-guard check models/ queries/` exits 1 on a missing table or column.`schema-guard snapshot --dbt target --check` exits 1 if the committed snapshot is out of date | \n| **Any agent** (AGENTS.md, .cursorrules) | paste [rules/schema-guard.md](https://github.com/idk-arsh/schema-guard/blob/master/rules/schema-guard.md) | \n\n| Source | Command | Needs | \n|---|---|---|\n| dbt | `--dbt target` | `dbt docs generate` (catalog.json). With only manifest.json, tables are checked but columns aren't | \n| DuckDB / SQLite | `--duckdb wh.duckdb` /`--sqlite app.db` | nothing | \n| Postgres, MySQL, Redshift ... | `--url postgresql://...` | `sqlalchemy` + driver | \n| BigQuery | `--bigquery my-project.my_dataset` (or`region-us` ) | the `bq` CLI; INFORMATION_SCHEMA queries are free | \n| Snowflake | `--snowflake MY_DB [--connection name]` | `snowflake-connector-python` ,`~/.snowflake/connections.toml` | \n| Databricks | `--databricks my_catalog` | `databricks-sql-connector` ,`DATABRICKS_HOST` /`_HTTP_PATH` /`_TOKEN` | \n| Migrations or a schema dump | `--ddl migrations/` | nothing; CREATE / ALTER / DROP applied in file order | \n| Anything else | `--csv columns.csv` | an export of `information_schema.columns` | \n\nSeveral files in `.schema-guard/` are merged, so one repo can cover more than one warehouse. The person taking the\nsnapshot needs warehouse access once; the agent never does.\n\nI ran the Databricks reader against a fresh Databricks Free Edition workspace to prove the Unity Catalog path works end to end.\n\n```\n$env:DATABRICKS_HOST = \"<workspace>.cloud.databricks.com\"\n$env:DATABRICKS_HTTP_PATH = \"/sql/1.0/warehouses/<id>\"\n$env:DATABRICKS_TOKEN = \"<token>\"\n\npython -m schema_guard.cli snapshot --databricks samples -o .schema-guard/databricks-samples.json\n# wrote .schema-guard\\databricks-samples.json: 9 tables, 277 columns, dialect databricks\n\npython -m schema_guard.cli check \"SELECT customerid, first_name FROM samples.bakehouse.sales_customers LIMIT 10\"\n# (silent: passes)\n\npython -m schema_guard.cli check \"SELECT customer_id, first_name FROM samples.bakehouse.sales_customers LIMIT 10\"\n# <sql>: `bakehouse.sales_customers` has no column `customer_id`. Did you mean: `customerid`?\n```\n\nThe snapshot came back in under 30 seconds on a cold warehouse. The reader pulls from `information_schema.columns`,\nso any Unity Catalog you can read works the same way.\n\n- **Shell commands:**`bq query` ,`snowsql` ,`snow sql` ,`psql` ,`duckdb` ,`sqlite3` ,`databricks` ,`spark-sql` ,`mysql` ,`trino` , SQL passed to scripts (`python run_sql.py \"...\"` ,`python -c \"...sql...\"` ), heredocs and`-f file.sql` .\n- **Files:** every`.sql` the agent writes or edits. dbt`{{ ref() }}` and`{{ source() }}` are resolved to real\ntables. On an edit, only problems the edit*adds* are reported, so old debt in a file doesn't block new work.\n- **MCP tools:** any tool with a`sql` /`query` /`statement` argument (Snowflake, Databricks, BigQuery, Postgres\nMCP servers).\n- **Resolution:** CTEs, subqueries, correlated subqueries, aliases,`USING` , set operations, CTAS and temp tables\ncreated earlier in the same script, INSERT column lists, UPDATE SET, DELETE WHERE. Parsing is by[sqlglot](https://github.com/tobymao/sqlglot) , so 20+ dialects.\n\nA false block costs more trust than a missed one, so it says nothing when it can't be sure:\n\n- SQL it can't parse, Jinja beyond ref/source/config, sources it can't see into (UNNEST, LATERAL, table\nfunctions, PIVOT), `SELECT *` from a table it doesn't know, struct and JSON field access.\n- Tables from a database the snapshot doesn't cover (unless the name is a near miss of one it does).\n- **Stale snapshot:** if the agent sends the exact same SQL again after a denial, it goes through. A new column can\nslow the agent down once but never lock it out. Refresh with`schema-guard snapshot` , and put`--check` in CI.\n\nA guard that blocks valid SQL gets uninstalled, so this matters more than the catch rate.\n\n| Corpus | Valid queries | False blocks | Planted wrong names caught | \n|---|---|---|---|\n| [Spider](https://yale-lily.github.io/spider) dev, 20 databases (held out: never looked at while building) | 1,034 | **0** | 1,032 / 1,034 | \n| [defog sql-eval](https://github.com/defog-ai/sql-eval) , 7 databases × Postgres, BigQuery, Snowflake, MySQL, SQLite | 960 | 0 | 959 / 960 | \n\nEvery valid query is human-written gold SQL that runs on its database, so any finding would be a false block.\nThe planted mistakes swap one real name for a wrong one the way agents get it wrong (a column from another table,\n`_id` / plural / `_name` variants, singular vs plural table names). I fixed 3 checker bugs that defog exposed, so\ntreat its numbers as training numbers; Spider is the honest one. Its 2 misses are inside correlated subqueries, where\nthe checker deliberately gives the benefit of the doubt. Run them: `python evals/benchmark_spider.py`,\n`python evals/benchmark_defog.py` (needs `pip install defog-data`).\n\n- It checks names, not meaning. `SUM(gross_amount)` when you wanted`net_amount` passes.\n- Dynamic SQL built from string pieces in application code isn't seen.\n- The snapshot is only as fresh as the last `schema-guard snapshot` .\n- The Snowflake and BigQuery snapshot readers are unit-tested on their output format, not yet run against live\naccounts. The Databricks reader was run live on Free Edition on 2026-10-04 (see below). If you run one,\n[tell me how it went](https://github.com/idk-arsh/schema-guard/issues) .\n\nPart of a set of small, measured tools for AI agents working on data:\n[data-agent-rules](https://github.com/idk-arsh/data-agent-rules) (rules + safety hooks, cost checks, masked previews),\n[show-your-sql](https://github.com/idk-arsh/show-your-sql) (every number in the answer traced to a query result),\n[data-test-guard](https://github.com/idk-arsh/data-test-guard) (agents can't delete or loosen tests to go green).\n\nMIT licensed.\n[Using it? Open a](https://m8ven.ai/mcp/idk-arsh/schema-guard?s=readme) [\"We use this\" issue](https://github.com/idk-arsh/schema-guard/issues/new?template=we-use-this.yml) or add a line to [ADOPTERS.md](https://github.com/idk-arsh/schema-guard/blob/master/ADOPTERS.md). False blocks are the bug I most want to hear about.", "url": "https://wpnews.pro/news/schema-guard-stop-ai-agents-from-inventing-column-names-in-sql", "canonical_source": "https://github.com/idk-arsh/schema-guard", "published_at": "2026-10-06 00:46:04+00:00", "updated_at": "2026-10-06 01:18:45.858027+00:00", "lang": "en", "topics": ["ai-agents", "ai-tools", "developer-tools", "ai-safety"], "entities": ["schema-guard", "Snowflake", "Claude Haiku 4.5", "Claude Sonnet 5", "Claude Code", "DuckDB", "GitHub"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/schema-guard-stop-ai-agents-from-inventing-column-names-in-sql", "markdown": "https://wpnews.pro/news/schema-guard-stop-ai-agents-from-inventing-column-names-in-sql.md", "text": "https://wpnews.pro/news/schema-guard-stop-ai-agents-from-inventing-column-names-in-sql.txt", "jsonld": "https://wpnews.pro/news/schema-guard-stop-ai-agents-from-inventing-column-names-in-sql.jsonld"}}