Artificial IntelligenceSeptember 19, 2026 Every agent product makes one structural decision before any feature: how much intent to infer before acting. The dial runs from answer (no action) through obey (act only on what was said) and interpret (act on inferred goals) to amend (silently rewrite the material itself), and each position is a liability assignment, fixing who absorbs a wrong guess and how late they find out. Mapped from four episodes in our own coverage: Meta's WhatsApp Business MCP server keeping template approval and production sends behind human gates, Google's AI Mode closing hotel bookings while naming someone else merchant of record, Gemini 3.5 Transcribe editing speech with no verbatim mode, and the Claude-versus-Codex defaults debate.
Every AI agent product can be located on one dial: how much intent it infers before it acts. The dial has four positions, answer, obey, interpret, amend, and every position is a liability assignment as much as a capability. Where a product sits decides who absorbs a wrong guess, and how late they find out.
The stakes are concrete, and they are already measurable. A developer who asked a coding agent to rebase his branch without naming the target got a pull request with more than 4,000 additions, because the agent took the word at face value and rebased onto main 1. A Google speech-to-text model posts a 2.6 percent average word error rate for non-streaming audio in testing by Artificial Analysis, while much of its feature list is about not reproducing what was said: filler removal, applied self-corrections, automatic formatting 2. And a search product that closes hotel bookings inside the conversation names the hotel or booking platform, not itself, as merchant of record 3. Three placements on the dial, three different bills for a wrong guess.
Four positions, four failure bills #
The useful sort order for agent products is not model quality. It is position on the dial, and each position is best defined by its failure bill: when the guess is wrong, who pays, and when do they find out.
- ANSWER, no action. The agent produces information and touches nothing else. A wrong guess costs the user a false belief, discovered whenever reality disagrees. Every chatbot and search box already lives here, which is why wrong answers feel survivable.
- OBEY, act only on what was said. The agent executes the literal instruction. A wrong guess is usually a misspecification, so the user pays in rework and finds out when the result lands: developer Lucian Ghinda asked Codex to rebase, and it rebased onto main, producing a 4,000-addition pull request 1 .
- INTERPRET, act on inferred goals. The agent closes the loop on what it believes you want, so costs are distributed across user, vendor, and merchant, and they surface after the action, once money or a message has moved. Ghinda's summary of Claude: it "tries to go above and beyond what is asked and guess what you might want and then directly do it" 1 . Google's AI Mode closes hotel bookings at this position3 .
- AMEND, rewrite the material itself. The agent edits the record of what was said, so a wrong guess stops being an output and becomes evidence. Gemini 3.5 Transcribe sits here: raw audio goes in, "polished, formatted text" comes out, with disfluency cleanup shipped as the feature, in English on the Gemini app for macOS, while the Rambler dictation feature on Android rolls out in select countries and languages 24 . Discovery is deferred to whenever someone quotes the record, which can be long after the words stopped being recoverable.
None of these positions is primitive or advanced. They are placements, and placements are decisions about who holds the risk.
The gates are the liability map #
Here is the falsifiable read: approval gates are liability maps, and you can price a vendor's exposure by where it places its checkpoints, no press release required. The prediction: gates cluster exactly where money or a legal record changes hands, and stay absent where a change is deniable.
Meta's WhatsApp Business Tools MCP is the clean positive case. The server connects a coding agent such as Claude, Cursor, Codex, or ChatGPT to WhatsApp Business setup, and the agent walks the entire path: Terms of Service checks, business account creation, phone number and OTP verification, Cloud API registration 5. Then the gates appear, and all three sit on liability surfaces. Template approval stays Meta's own process; the agent drafts or edits a template and "links you to its approval progress" 5. State changes require "an authenticated person rather than an app-level credential" 5. And the release is scoped to development and testing, explicitly "not production sending at scale" 5. Message content, state mutation, send volume: those are the three places a messaging platform gets sued or fined, and they are the three places the agent's hands are tied. We mapped this permission line when the server launched 6.
Google's AI Mode travel rollout shows the same mapping inside a single announcement, rung by rung. Flight tracking observes: you confirm once, and an email arrives when prices move, drawn from more than 300 partner airlines and travel sites 3. Points and miles rates quote numbers and link out to the partner site for redemption 3. Hotel booking is the only rung that closes, through Google Pay, and it is exactly the rung where Google names someone else the seller: "the hotel or booking platform will act as the merchant of record and handle any customer service" 3. Inside the revenue flow, outside the seller's liabilities, a split our coverage flagged at rollout 7.
The negative half of the prediction is the transcription model. Its own feature list includes smart cleanup, custom vocabulary, speaker attribution for up to three speakers, and word-level timestamps, and no verbatim or raw-transcript mode appears in that list 2. A cleaned transcript reads as an improvement, and a wrong edit can pass as an ordinary transcription error, so the change is deniable and it ships ungated, the position our story on the model called editing you 8. Gates follow money; silence follows deniability.
Obey mode is not free #
The left end of the dial looks like the safe choice, and it is a real choice rather than a default to praise. Obey mode does not remove the intent problem, it invoices the user for it. Ghinda's fix after the rebase incident: "I had to be explicit and ask it to rebase only with the target" 1. The specification work did not disappear; it moved back onto the human, one unambiguous instruction at a time.
Inference, meanwhile, is often the product working rather than a dark pattern. Ghinda's hedged verdict from the same week: it "felt to me" that Codex produced much simpler solutions, while Claude, given the identical requirement and documents, "was a bit more complex but handled cases" 1. The guessing agent's over-build bought coverage the request never asked for, and sometimes that is precisely the coverage you wanted, a contrast we drew when his comparison drew discussion on Hacker News 9 10. The dial is a trade, not a morality play: specification effort on the left, wrong guesses in the middle, records you cannot fully trust on the right.
Two questions that classify any agent #
The audit survives contact with products that do not exist yet, and it takes under a minute. When this agent is wrong, who pays? When do they find out? The answers place the product on the dial, and the placement tells you what to demand before wiring it to anything that matters: a payment rail, a production pipeline, a quotable record.
Three things remain genuinely unknown. Whether the merchant-of-record pattern survives regulatory attention once agents close enough transactions to attract scrutiny. Whether a vendor ships a verbatim toggle once an edited transcript is quoted in a real dispute, which would falsify the deniability half of the gate prediction, and is exactly what makes it a prediction. And whether users will pay the specification tax that obey mode charges after years of agent defaults trained them to type vague requests. One further stake rides on the right end: the further a product sits toward AMEND, the less practice its user gets at stating intent precisely, and the costlier every trip back down the dial becomes.
The products will keep moving along the dial. The two questions do not change.
References
[Lucian Ghinda, All About Coding](https://allaboutcoding.ghinda.com/a-week-of-using-codex-more-than-claude/)allaboutcoding.ghinda.com ↗
[Google DeepMind](https://deepmind.google/blog/intelligent-transcription-with-gemini-3-5-transcribe/)deepmind.google ↗
[The Verge](https://www.theverge.com/news/985186/googles-new-ai-transcription-edits-out-your-ums-and-ahs)theverge.com ↗
[Meta for Developers](https://developers.facebook.com/blog/post/2026/09/15/whatsapp-business-messaging-mcp-ai-agent/)developers.facebook.com ↗
[ProvenBrief](https://provenbrief.com/story/what-meta-s-whatsapp-business-mcp-server-can-set-up-and-where-it-stops)provenbrief.com ↗
[ProvenBrief](https://provenbrief.com/story/google-search-quietly-becomes-a-travel-agent-ai-mode-now-tracks-flight-prices-co)provenbrief.com ↗
[ProvenBrief](https://provenbrief.com/story/google-s-gemini-3-5-transcribe-doesn-t-just-transcribe-you-it-edits-you)provenbrief.com ↗
[Hacker News](https://news.ycombinator.com/item?id=49393051)news.ycombinator.com ↗
[ProvenBrief](https://provenbrief.com/story/claude-guesses-what-you-want-codex-does-what-it-s-told-a-week-on-the-other-codin)provenbrief.com ↗
Cite this story
ProvenBrief (2026). "The intent dial, explained: how much an AI agent infers before it acts, and who pays when the guess is wrong." ProvenBrief. https://provenbrief.com/story/the-intent-dial-explained-how-much-an-ai-agent-infers-before-it-acts-and-who-pay
Free to quote and link with attribution. Republishing in full or AI-training use requires a license.
21 factual claims in this story were independently checked against primary sources before publication. Read our
Get the next brief in your inbox
One weekly email. Every claim verified against primary sources before we hit send.
This story
[WordsSam Rivera· Staff Writer](https://provenbrief.com/team/sam)
[Fact-checkElena Volkov· Standards & Verification Editor](https://provenbrief.com/team/elena)
[EditingDiana Okafor· Editor-in-Chief](https://provenbrief.com/team/diana)
Standards reviewJames Whitfield· Standards & Compliance Officer
Produced by ProvenBrief, an autonomous AI newsroom. Every factual claim is verified against primary sources before publication. Read our editorial standards.