Same prompt, same endpoint, different score
A test of a small prompt edit on 100 single-call tool-calling cases from UC Berkeley's BFCL V4 benchmark found the edit reduced invented optional arguments from 14 to 7 across three runs but produced …
A test of a small prompt edit on 100 single-call tool-calling cases from UC Berkeley's BFCL V4 benchmark found the edit reduced invented optional arguments from 14 to 7 across three runs but produced …
OpenRouter's monitoring of the deepseek/deepseek-v4-flash-0731 endpoint at mancer/fp8 recorded a B3IT level shift of 0.36 total-variation distance, while the LT logprob probe failed because the endpoi…
OpenRouter has tracked the deepseek/deepseek-v4-flash-0731 endpoint at streamlake/fp8 since 2026-08-13, recording $0.032 in total spend across 66,386 queries. The B3IT border-input detector measured a…
OpenRouter retired monitoring of the deepseek/deepseek-v4-flash-0731 endpoint at akashml/fp8 on 2026-09-11 after it stopped yielding usable samples, leaving its B3IT tracking status retired and stalle…
Whatstrending.ai, a service that tracks daily request counts for AI models via OpenRouter's usage feed, reports that OpenAI's flagship model gpt-5.6-luna experienced a 38.3% drop in daily usage over t…