Schema-guard, stop AI agents from inventing column names in SQL Schema-guard, an open-source tool that snapshots real table and column names into a repo file (.schema-guard/schema.json) and validates AI coding agents' SQL against it before the SQL runs, produced 48 of 48 runnable files across 4 requests × 3 arms × 3 runs on Claude Code, versus 0 of 24 without the snapshot, according to the project's published eval. In the eval, Claude Haiku 4.5 and Sonnet 5 both failed all 12 baseline files on missing column names, while the snapshot-plus-hook arm ran 12 of 12 files for each model, with 12 of 12 correct for Haiku and 11 of 12 for Sonnet. The hook arm cost about the same as baseline ($0.65 vs $0.58 for 12 Haiku runs; $1.61 vs $1.61 for Sonnet), and the author discloses the eval is small and synthetic with 3 runs per cell and two grader references added after reading runs. Your AI agent stops inventing column names. Coding agents write SQL against the schema they think you have, based on your README, an old query, or a naming convention. Then it fails in CI, in a dashboard, or at 2am. Snowflake's own developer blog ran a whole post on this in September 2026 My coding agent won't stop hallucinating table columns https://www.snowflake.com/en/developers/blog/coding-agent-hallucinating-table-columns/ . schema-guard keeps a snapshot of your real tables and columns in the repo names and types only, no data, no credentials and checks the agent's SQL against it before it runs or lands in a file : schema-guard: this SQL names things that are not in the schema snapshot .schema-guard/schema.json, taken 2026-09-29T01:35:07Z : - analytics.customers has no column country . Did you mean: country iso2 ? - analytics.orders has no column customer id . Did you mean: cust id , order id ? - analytics.customers has no column id . Did you mean: cust id ? Fix the names and try again. ... That's a real denial from the eval below: Claude Haiku 4.5 writing models/revenue by country.sql from a README that describes last year's schema. The agent reads the denial, fixes its SQL and moves on. You never see the broken version. Setup. An analytics repo whose README describes an older schema customer id , created at , country , plus two up-to-date models that use a few of the real names. The agent can read and write files but can't reach the warehouse. That's the situation in Snowflake's post: the agent has the repo, not the account. Each request asks for a new SQL model, and afterwards the grader runs every file the agent wrote against the real DuckDB warehouse. 4 requests × 3 arms × 3 runs, on Claude Code. | Arm | Haiku 4.5: fails on a missing name / runs / correct | Sonnet 5: fails / runs / correct | |---|---|---| | baseline repo only | 12 / 0 / 0 of 12 | 12 / 0 / 0 of 12 | | rule snapshot + one line in CLAUDE.md | 0 / 12 / 11 of 12 | 0 / 12 / 12 of 12 | | hook snapshot + hook, no instruction | 0 / 12 / 12 of 12 | 0 / 12 / 11 of 12 | What that means: - Without a snapshot, neither model wrote one working file 0 of 24 . Both trusted the README, and even when they copied real names from the existing models they mixed them with stale ones. - With the snapshot, every file ran 48 of 48 . The 2 wrong answers are logic errors, not names: Haiku started weeks on Sunday, and Sonnet counted the last days of 2025 in the first week. - The hook is the safety net for agents that don't go looking. Haiku was denied in 10 of its 12 hook runs and fixed the names on the first retry every time. Sonnet found .schema-guard/schema.json by itself and was never denied. A one-line rule gets the same result if the agent follows it; the hook doesn't depend on that, and it also covers ad-hoc queries and MCP tools. - Cost: the hook arm cost about the same as baseline $0.65 vs $0.58 for 12 Haiku runs; $1.61 vs $1.61 for Sonnet . Be skeptical of this: the world is small and synthetic, and the stale README is designed in docs drift is normal, but I chose how far . There are 3 runs per cell. Two grader references were added after I read runs: "net revenue" net of refunds it changed 2 Haiku grades, one rule run and one hook run , and listing all 52 weeks with zeros 5 Sonnet grades . Both are disclosed in evals/scenarios.py https://github.com/idk-arsh/schema-guard/blob/master/evals/scenarios.py , and every run's SQL is in evals/results/ https://github.com/idk-arsh/schema-guard/blob/master/evals/results . Rerun it: cd evals && python run eval.py --model