Cline in Production: BYO-Key Costs, MCP Limits, and Terminal-Bench Results A team evaluating the Cline VS Code agent extension found that a BYO-key setup with Kimi K3 on Terminal-Bench 2.1 raised its pass rate from 77.5% (69/89) to 88.8% (79/89) while cutting run cost from $79 to $49.80, attributing the savings to fewer sessions dying in retries, loops and self-termination. The team cautions that the benchmark is better treated as a harness-debugging case study than a definitive ranking, that a separate Python MCP SDK test failed during initialization without establishing a Cline client defect, and that Cline should be treated as an agent harness with an editor front end rather than a security boundary. We did not test Cline because we needed another autocomplete tool. We tested it because flat per-seat AI IDE pricing makes cost attribution difficult, while closed agent runtimes make model and tool migrations expensive. Cline + Kimi K3 on Terminal-Bench 2.1: score up, spend downBaseline pass rate 69/89 77.5%/100Confirmation pass rate 79/89 88.8%/100Baseline run cost $79Confirmation run cost $49.80 The higher-scoring run on the same 89-task suite was also the cheaper one, because fewer sessions died in retries, loops, and self-termination. Cline separates the agent runtime from the inference provider. The VS Code extension supplies the agent loop, file operations, terminal integration, approvals, context management, and MCP client. We supply the model account and pay the inference provider directly. That separation gives us three things we cannot assume from a bundled IDE subscription: That does not make Cline free. It exchanges predictable seat pricing for variable inference spend, provider rate limits, API-key management, and substantially more operational responsibility. We evaluated four practical questions: The short answer is mixed. BYO-key model choice and usage-based accounting provide useful control, but we did not verify the extension's installation or provider-configuration workflow. Our separate Python MCP SDK test failed during initialization; it does not establish a Cline client defect. The benchmark is much more useful as a harness-debugging case study than as a definitive “Cline beats Cursor” ranking. Cline’s SDK production architecture also matters. We can instrument model-call, tool-use, session-end, and usage events, including token counts and finish reasons. The SDK supports iteration limits, per-turn token limits, loop detection, mistake limits, and host-side cancellation. Those are the controls we expect from an agent runtime. They are not proof that a developer workstation is safely sandboxed. Our evaluation therefore treated Cline as an agent harness with an editor front end, not as a security boundary. For a team rollout, we would verify the current Marketplace extension identifier and test installation in a clean VS Code profile. The commands below are illustrative; we have not verified their identifier or execution: code --install-extension saoudrizwan.claude-dev code --list-extensions --show-versions | grep -i saoudrizwan.claude-dev We would confirm the publisher, display name, and exact extension identifier against the current Marketplace listing before using these commands. For a team rollout, we would pin and test a known version before broad deployment: code --install-extension saoudrizwan.claude-dev@