Bedrock and LangChain disagree on what input_tokens means — never price it without knowing who counted A developer building a usage ledger for an AI copilot inside a regulated SaaS product found that Amazon Bedrock's raw InvokeModel API and LangChain's ChatBedrockConverse report the same input_tokens field under different conventions: raw InvokeModel reports the uncached remainder while LangChain's usage_metadata already includes cache reads and writes, so a pricing function that treats them alike double-bills the cached portion of every prompt. The developer's fix makes the counting convention a required argument of the price calculation, and the accompanying test suite shows the same token counts priced at $0.000545 under the cache-folded convention versus $0.001445 under the additive one. I was adding cache-read and cache-write columns to a usage ledger, and I put two usage payloads side by side to get the parsing right. Both came from Claude on Bedrock. Both had input tokens and both reported prompt-cache activity. In one, input tokens was what was left after the cache. In the other, it already included the cache. So: same model family, same AWS account, same field name. A pricing function that treats them alike bills the cached part of every prompt twice. And prompt caching is the feature you turned on to make the bill smaller . Here's the claim I want to defend: 💡 Ask who counted. A token count only becomes a number once you know the convention it was reported under. Make that convention a required argument of the price, never an assumption buried inside it. InvokeModel with an Anthropic messages body reports input tokens as the usage metadata , as our client reads it, reports the This is an AI copilot inside a regulated SaaS product. Every model call lands in a usage ledger and gets priced against a versioned rate card. Two stacks feed the ledger. One is a .NET feature that turns natural-language questions into search filters and calls InvokeModel directly. The other is a Python stack, the document-search copilot plus a question-generation feature, that goes through LangChain's ChatBedrockConverse . | | Raw InvokeModel Anthropic messages body | LangChain usage metadata via ChatBedrockConverse | |---|---|---| | input tokens means | the uncached remainder | the full input, cache included | | Cache read | usage.cache read input tokens | input token details.cache read | | Cache write | usage.cache creation input tokens | input token details.cache creation , or ephemeral 5m input tokens + ephemeral 1h input tokens | | Total input is | input + read + write | input | | Cache slices are | additive | a breakdown | The Python client, at the point where it stopped throwing the breakdown away: details = raw.usage metadata.get "input token details" or {} A per-TTL cacheDetails breakdown puts the write under the ephemeral keys and may leave cache creation at zero, so accept either representation. ephemeral = int details.get "ephemeral 5m input tokens", 0 or 0 + int details.get "ephemeral 1h input tokens", 0 or 0 usage = { "input tokens": raw.usage metadata.get "input tokens", 0 , already includes the slices below "cache read tokens": int details.get "cache read", 0 or 0 , "cache write tokens": int details.get "cache creation", 0 or 0 or ephemeral, } That's a second gotcha in the same dict. If you read only cache creation , your cache writes can quietly disappear. The .NET side has its own: JsonElement.TryGetInt32 throws on a non-numeric element, so the parser checks ValueKind first. Otherwise one malformed count fails the whole request. On scope: I'm vouching for raw InvokeModel with the Anthropic body, and for LangChain's usage metadata as our client receives it. This code doesn't tell me whether the folding happens in Converse or in langchain-aws, and I haven't checked other Bedrock APIs. I'll stick to the test suite's numbers here. They're priced at its Sonnet 5 fixture rates the rate card's original list-price seed : $2 input, $10 output, $0.20 cache read and $2.50 cache write per million tokens. Fact public void Price Takes The Cache Slices Out Of The Input Count On A LangChain Row { // Same numbers reported the other way: the 520 already contains the 400 read and the 50 written, so // only 70 tokens are charged at input rate. Charging all 520 would bill the cached prompt twice. var priced = PricingEngine.Price Call TokenConvention.CacheFolded, input: 520, output: 20, cacheRead: 400, cacheWrite: 50 , Sonnet5 ; Assert.Equal 0.000545m, priced.Usd ; Assert.Equal 70, priced.BillableInputTokens ; } Its additive twin, with the same counts, asserts $0.001445. The end-to-end fixture is a cached copilot turn: 12,000 input tokens, of which 10,000 were read from cache and 1,500 written, plus 800 output. Priced correctly it's $0.01475. Through the wrong branch it's $0.03775, which is 2.56×. The multiplier grows with the cache-hit rate, too. The better your caching works, the worse the overbill. The fix isn't clever. A metered call can't be constructed without saying how it counts: public enum TokenConvention { /// Raw Bedrock: the input count is the uncached remainder, so the cache slices add to it. CacheAdditive, /// LangChain: the input count already includes the cache slices, so they break it down. CacheFolded, } // on MeteredCall public required TokenConvention Convention { get; init; } required means there's no default to fall back on. Ledger rows keep the counts exactly as reported, and in the PR I told reviewers the column comments stating each convention are the contract. Later, when usage moved to events pushed through an ingest endpoint, an event that doesn't declare cache additive or cache folded started getting rejected instead of defaulted. Ask who counted, and ask at the door. On a folded row the engine subtracts the slices and floors at zero, so an inconsistent row can't become a negative charge. This is the version I merged: // The cache slices are always priced at their own rates. What differs is whether the input count // still contains them: under CacheFolded it does, so subtract them out before charging input rate. var billableInput = call.InputTokens; if call.Convention == TokenConvention.CacheFolded { var uncached = decimal call.InputTokens - call.CacheReadTokens - call.CacheWriteTokens; billableInput = uncached <= 0 ? 0 : long uncached; } var usd = PerMillion billableInput, rate.InputUsdPerMTok + PerMillion call.OutputTokens, rate.OutputUsdPerMTok + PerMillion call.CacheReadTokens, rate.CacheReadUsdPerMTok + PerMillion call.CacheWriteTokens, rate.CacheWriteUsdPerMTok ; static decimal PerMillion long tokens, decimal? usdPerMTok = tokens usdPerMTok ?? 0m / 1 000 000m; Read that first comment again. Always. A chat row on the rate card is allowed to leave out its cache rates. It has to be, because gpt-oss has no prompt cache. Price a folded row against one of those, and 450 of the 520 input tokens come out of the input count and then get priced at ?? 0m . They're removed from what the input rate charges, and they're charged nothing on their own. On that fixture you get $0.00034 instead of $0.00124. I missed it. GitHub Copilot's reviewer didn't, at 17:03:08 UTC: "In CacheFolded mode, billable input always subtracts cache read/write tokens. If the resolved rate exists but has null cache slice rates allowed for chat rows, e.g. models without prompt cache , those cache tokens will be subtracted out of input and also priced at $0 because PerMillion ..., null = 0 , effectively making them free. Subtract only the slices that will actually be billed separately." I merged at 17:03:58, and the comment went in unaddressed. This wasn't a billing incident. Nothing called the engine yet, because the rollup that would consume it hadn't landed, so the bug never priced a real bill. I also don't want to make the bot the hero or the punchline. It read the null-rate case more carefully than I did, and I merged faster than I read. The fix was committed under two minutes after the comment and landed in a follow-up PR the next day: js var pricedSeparately = rate.CacheReadUsdPerMTok = null ? decimal call.CacheReadTokens : 0m + rate.CacheWriteUsdPerMTok = null ? decimal call.CacheWriteTokens : 0m ; var uncached = call.InputTokens - pricedSeparately; A slice only leaves the input count if something else is about to charge it: Fact public void Price Charges A Folded Cache Slice As Input When No Rate Covers It { var noCacheRates = Rate "m1", input: 2.00m, output: 10.00m ; var priced = PricingEngine.Price Call TokenConvention.CacheFolded, input: 520, output: 20, cacheRead: 400, cacheWrite: 50, modelId: "m1" , noCacheRates ; // All 520 at input rate, not 70, and the 450 cached tokens are not lost. Assert.Equal 520, priced.BillableInputTokens ; Assert.Equal 0.00124m, priced.Usd ; } Charging an unrated read at the input rate overstates it, since in the seed a read costs a tenth of input. I can live with that. The commit message says it better than I can now: it "over-states a cache read against its real discount but never loses a token that was sent." The two errors aren't symmetric. An overstated read is bounded, and it shows up as a ledger running hot against the invoice. A lost token shows up nowhere. That same rule drove the unglamorous half of the work. Fail paths now keep tokens that were already spent. Before, two error paths empty response, unparseable response never wrote the request row at all, which left it in processing forever with its tokens gone. The rate card already rejects the shapes it can prove wrong. A chat row without an output rate would bill every answer for free, so it's refused. This shape can't be refused. As I put it in review: "Some models have no prompt cache and legitimately carry none gpt-oss , so requiring them would reject a correct row; a model that does cache and was seeded without them underbills its cached slices. Reconciling the ledger against the AWS invoice is what catches that, not the schema." The same reviewer found two more things earlier. One was long arithmetic in that subtraction, which could wrap a corrupted row into "an extreme overcharge instead of being floored at 0". It's decimal now. The other was integration tests that did a plain return when there was no database, which counts as a pass. They were green having asserted nothing. Once they skipped explicitly, the no-database run went from 53 skipped to 60. That first bullet is the strongest argument against my own design. With two conventions in the pricer, every edge case has to be right in two places. The usage rollup I wrote later already normalizes input to the total for its token sums, so the alternative is sitting in my own code: normalize once at ingestion and keep the pricer single-path. I carried the convention all the way to the price instead, so every ledger row still holds exactly what its provider reported. The reason is repricing: a past window gets repriced by running the engine again over the raw rows. If I ever file a source under the wrong convention, I fix the label and recompute, and the evidence is still in the ledger. A row normalized at the door has already thrown that evidence away. So would you normalize at the boundary and price a single shape, or carry the convention through to the price and own two branches? Appreciate you reading an entire post about subtraction 🧮 If you meter prompt caching through more than one SDK, or your ledger has ever drifted from the AWS invoice by a suspicious amount, tell me which convention bit you first. I'm on LinkedIn https://www.linkedin.com/in/rodrigo-diego-67867185/ .