I Tracked Every Cost-Cutting Trick AI Teams Used This Week. Here’s What Actually Works Spotify Engineering Principal Product Manager Dimitri Mazmanov reported roughly 90% mean token savings across four scenarios on a Java monorepo by using a PreToolUse hook that blocks Claude Code file reads over 350 lines (configurable via the SHUNT_MIN_LINES environment variable) and delegates bulk reads to a cheaper worker model, Gemini 2.5 Flash, which returns structured bullets. Mazmanov noted the savings apply only to bulk-read work, that delegation adds 10 to 30 seconds of latency (capped at 30), and that the worker missed a subtle thread-safety bug Claude caught in seconds, so editing, debugging and safety-critical analysis stay on the frontier model. Last week I kept a running list of every concrete cost-cutting technique AI teams shared in public. By Saturday it had 191 items on it. Most of them were noise: a “we saved 60%” with no method behind it, a dashboard screenshot, a thread that ended in a link to a course. Six of them were different. Each one named a mechanism, gave a number, and the good ones admitted where it breaks. I’m going to walk through those six, then pull out five principles I think a .NET/AI developer can apply this week, then spend a few paragraphs on a study that made me distrust my own excitement. One rule for this article. Every number is labeled as reported by the source, computed by me from the source’s numbers, or illustrative. Where I could not open a primary source, I say so. Some of this week’s most shared figures were hypothetical in the original posts, and I would rather show you a smaller honest list than a big shaky one. The most shared item of the week came from Dimitri Mazmanov, a Principal Product Manager at Spotify Engineering, writing about a tool called Portal. His problem is one most of us have: Claude Code reads huge files with the expensive model, and most of those reads are just “find me the thing”. His fix has three layers. First, a PreToolUse hook blocks any file read over 350 lines configurable through an environment variable, SHUNT MIN LINES and tells the agent to delegate instead. A second hook catches the workaround where the agent tries cat or head on the same file through bash. Second, wrapper scripts send the bulk read to a cheaper worker model, Gemini 2.5 Flash in his setup, which returns structured bullets. Third, a skill file documents when to call which script. What I liked most was one line from the write-up: prompt instructions are suggestions, hooks are architecture. The agent does not get a vote. Reported result: about 90% mean token savings across four scenarios on a Java monorepo. Two caveats that matter. First, that is savings on the bulk-read work, not a 90% cut of the whole bill, and the worker’s tokens still cost money, just less. Second, delegation takes 10 to 30 seconds, capped at 30, so you are trading latency for cost. Then the part that made me trust the post. Mazmanov wrote that the worker found surface-level patterns but missed a subtle thread-safety bug, and Claude spotted it in seconds. So editing decisions, debugging, and safety-critical analysis stay on the frontier model. He also keeps edits there because the worker’s summaries lacked reliable line numbers. Here is my version of the hook. It is a sketch of the pattern, not Spotify’s code. bash /usr/bin/env bash .claude/hooks/check-file-size.shMAX LINES="${SHUNT MIN LINES:-350}"input=$ cat file=$ echo "$input" | jq -r '.tool input.file path // empty' if -n "$file" && -f "$file" ; then lines=$ wc -l < "$file" if "$lines" -gt "$MAX LINES" ; then echo "Blocked: $file has $lines lines. Run ./scripts/bulk-read.sh \"