The Model Is Rented. The Policy Is the Product. SpaceX's reported all-stock acquisition of Anysphere, the maker of Cursor, at $60 billion values a coding product whose real asset is the per-turn policy — what the model sees, which tools it can reach, and what counts as done — not the underlying model, which Cursor rents from Anthropic, OpenAI, Google, and xAI while also shipping its own Composer and Grok models. The author argues constraints enforced outside the prompt beat constraints written into it: tool returns need observed, unverified, and next-step fields, and dangerous actions like delete must be made unreachable rather than merely forbidden in prose. The piece cites a support agent that told a customer three times to replace a working API key because a failed credential check rendered an actionable string into context, and notes prompt length has become a correctness variable as chained steps past about three should become stateless transforms with tests. A $60 billion headline says the asset is the assistant, but it is not. The model is rented. What a coding product is worth is the policy on each turn: what the model is allowed to see, which actions are reachable, and what counts as done. The system-prompt paragraph is the readable slice of that policy. It is not what the price buys. The number people reach for is SpaceX’s reported all-stock acquisition of Anysphere, the maker of Cursor, at $60 billion https://www.reuters.com/legal/transactional/spacex-buy-anysphere-60-billion-2026-06-16/ . I understand the reflex. When a code editor gets that kind of valuation, the obvious next question is whether someone built a secret paragraph. But Cursor is not a thin wrapper on someone else’s model. It is an editor that calls Anthropic, OpenAI, Google, and xAI models, and it also ships its own Composer and Grok models in a first-party pool. “Same model as anyone” is only half true. Even so, two products on the same Claude, GPT, or Grok still behave like different species. The gap is not the prose. It is what each product stuffs into the context window before every turn and which tools it lets the model touch. I want to trace where that gap actually lives. Consider a support agent that tells a customer three times to replace a working API key. The failure starts not with the model, but with a tool error that fell through to a default next step. A check that could not verify the key rendered a string into the context, something like “could not verify the credential.” The model treated that string as an instruction and picked the replace path. Once the wrong path exists in the tool list, the model will walk it again. This is a structural problem, not a phrasing problem. Every tool return needs three fields: what was observed, what could not be verified, and what the next step is for. Attach a next step only when the check tested the thing that step is about. A real invalid API key should still get the replace path. A missing probe model should not. Before I rewrite a system prompt, I look at what the agent was given. If the tool output is a wall of text with no partition between finding and instruction, the prompt was never the weak point. The model just obeyed the only actionable text it could see. “Never delete user data” is a comment, not validation. If the delete action is still in the tool list, the model can call it. The guard has to make the dangerous state unreachable: verify first, then reveal the action. This is the part that matters. Constraints that live outside the prompt usually beat constraints written into it. A generated config should reject unknown keys so a model cannot smuggle in a temperature or model override. If adding ten rules starts dropping the first five, anything that must hold does not belong in the prompt. Prompt length has become a correctness variable. Take a mail agent as a worked spec. The boundary is small: a fixed sender, an explicit set of recipients, the allowed facts and actions, what needs a human, and how it checks the sent result. Then test the awkward case: the thread changed after the draft was written. If the system still sends, no wording pass will fix it. User text is data for the next instruction, not a second system prompt. One model call is actually carrying five jobs: instructions, state, verification, scope, and session handoff. Most harnesses blur them into one big prompt and then blame the model. I define done as tests passed, lint passed, a type-check passed, or a smoke run passed. Not “the response looked right.” Scope is one feature. A dark-mode request that also rewrites the CSS and starts a notification system has left the job. Past about three chained steps, I stop hoping it remembers. Each step becomes a stateless transform with the slice it needs, structured input and output, and a test. I have seen the night-and-day same-model story passed around as if it were a measurement; I keep it as an operator report. The measurement is whether the handoff survives a fresh session. Tool descriptions are written for one model’s priors. Swap the model, and the same words mean something else. I have started treating descriptions as API contracts and comparing the tool calls, not only the final answer. A harness that feels identical on every model is leaving capability unused. There is a known failure mode where combining a function tool with a JSON response format makes the model call a weather tool on every request and loop, as an OpenAI developer forum report https://community.openai.com/t/function-calling-looping-uncontrollably-and-calling-unnecessarily/931730 describes. That was not a bad model. That was a bad harness configuration. Routing can live outside the model call. Cursor’s own model page https://cursor.com/help/models-and-usage/available-models lists first-party Composer and Grok models next to frontier models from OpenAI, Anthropic, and Google, with an Auto router picking per request. A smaller model is enough until the knowledge is sparse and a person has to fill the gap. Detailed instructions do not fix that hole. Before I trust a rules file, I eval every line I add. The AGENTS.md benchmark https://arxiv.org/abs/2602.11988 is the cleanest version of that warning I have read: across agents and models, context files did not generally raise task success, and they raised inference cost by over 20% on average. That held for files a model wrote and for files a developer committed. Instructions in the file were followed. Repository overviews, the part vendors like to recommend, were not helpful. The paper’s own conclusion is the one I keep: these files are for non-standard practices, and any attempt to raise performance should be measured before you leave it in the repo. Operators report a cliff near five hundred lines, though I treat that as a heuristic to test, not a spec. Hyper-specific entries help: the exact error, the exact command, the production gotcha. Aspirational style lists and compliance checklists interfere. A file the model wrote for itself is a reported small negative next to one hand-written architecture line, like “game-specific logic lives in plugins.” Four context failures. Poisoning: a false line gets in and is treated as true. Distraction: extra text hides the point. Confusion: irrelevant text moves the answer. Clash: two conflicting lines produce an inconsistent one. Write, select, compress, delegate. Stale memory is worse than no memory. The agent ignores “save what you learned,” so capture has to be passive. The product system prompt is not CLAUDE.md or AGENTS.md. It decides how that file is read. That distinction gets lost. Operators report the injected copy can be truncated, and that a flag people assume replaces the system prompt only replaces the opening sentences unless a file override is actually applied. I have not found a primary doc that confirms the exact behavior, so I treat it as something to test. The safer pattern is to have the agent read the rules file with the read tool before acting. Resume from a notes file, not the full transcript. A line like “read lines 35–70 of notes.md” is cheap and bounded. The full transcript is not. Keep the always-on file small, and add the rest with a hook for the task that needs it. The strongest version of the opposing view: people copy leaked system prompts and assume they have the product. Steelman it fairly. The text does show habits. Keep going until resolved. Maintain a todo list. Name files this way, not that way. But the more useful thing in a leaked prompt is mechanical, not rhetorical. It shows what gets injected each turn: open files, cursor position, recent files, edit history, linter errors. It shows a tool list that avoids the destructive whole-file write. What it does not show is retrieval over the code tree, or the subagents, or the apply loop that turns an edit into a diff. Copying the paragraph does not give you the editor. And a long pasted prompt taxes the window in a way the real product never does. On the same repo, a short chat with the exact slice can beat an IDE that stuffs the project in. Those are different products, and they should stay different. Basic line completion is commoditized. What remains is who sees the whole tree. People who already pay for the model still pay for apply, tab, and the agent loop. Operators report that bringing your own key drops apply and composer; I hedge that unless current docs say otherwise. The one thing a Cursor staff note confirms https://forum.cursor.com/t/composer-2-5-fails-with-byok/164176 is that first-party Composer models always run on Cursor infrastructure and reject a custom key or proxy. Some people pay the editor only for tab. Tab is also reported to suggest code that ignores function signatures in the same file. Chat for planning and documents. Tab for the in-flow edit. An agent for the repo. Large-file edits fail as writes. Placeholders appear, or the output-token ceiling turns a full rewrite into a truncation. The move is to split into smaller modules, then edit. A tool that can see the running app moves more than another wording pass. The implementer does not inherit the designer’s reasoning, only a clean spec. Hard memory cap per role. A task-to-document index beats one giant file. Research traces stay out of the implementation context. Here is the challenge. Specify the world the model is allowed to see and touch. Then specify the world it must not. Polishing the sentence is the small part. The Model Is Rented. The Policy Is the Product. https://pub.towardsai.net/the-model-is-rented-the-policy-is-the-product-022d075a59b9 was originally published in Towards AI https://pub.towardsai.net on Medium, where people are continuing the conversation by highlighting and responding to this story.