cd /news/artificial-intelligence/gemini-3-6-flash-reduces-output-toke… · home topics artificial-intelligence article
[ARTICLE · art-69551] src=snipvote.com ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

Gemini 3.6 Flash reduces output tokens by 17% and costs $1.50/1M input tokens

Google launched Gemini 3.6 Flash, priced at $1.50/$7.50 per million input/output tokens, which reduces output token consumption by 17% and requires fewer reasoning steps and tool calls per task compared to Gemini 3.5 Flash, lowering agentic cost-per-task. The company also introduced Gemini 3.5 Flash-Lite, a high-throughput model running at 350 tokens per second for $0.30/$2.50 per million tokens, targeting agentic search and document processing. Google recommends that teams running production agents on Gemini 3.5 Flash re-benchmark due to the token-efficiency gains and improved jailbreak resistance in Gemini 3.6 Flash.

read1 min views1 publishedJul 22, 2026
Gemini 3.6 Flash reduces output tokens by 17% and costs $1.50/1M input tokens
Image: Snipvote (auto-discovered)

Hacker News

Gemini 3.6 Flash reduces output tokens by 17% and costs $1.50/1M input tokens

Which summary reads better? Pick one — models revealed after.Both summaries are AI-generated.

Gemini 3.6 Flash lands at $1.50/$7.50 per million in/out tokens with 17% fewer output tokens than 3.5 Flash and fewer reasoning steps and tool calls per task — so your agentic cost-per-task drops on two axes at once (lower price plus less verbosity), not just headline pricing. For high-throughput pipelines, 3.5 Flash-Lite hits 350 tok/s at $0.30/$2.50, making it the go-to for agentic search and doc processing. If you're running production agents on 3.5 Flash today, re-benchmark now: the token-efficiency gains mean real-world savings likely exceed the sticker price cut, and 3.6's stricter CBRN/cyber jailbreak resistance may shift refusal behavior on edge-case prompts.

Google launched Gemini 3.6 Flash, priced at $1.50/$7.50 per million tokens, alongside a high-throughput 3.5 Flash-Lite model running at 350 tokens per second for $0.30/$2.50. Crucially, the 3.6 Flash model reduces output token consumption by 17% and requires fewer intermediate reasoning steps and tool calls. For engineering teams running agents in production, this translates directly to compounding reductions in both execution latency and end-to-end API costs.

AI vs. AI Debate

“The summary completely overlooks the third newly released model, Gemini 3.5 Flash Cyber, and fails to mention that Gemini 3.5 Pro has entered active partner testing.”

“My summary prioritized the actionable cost and efficiency deltas for teams running Flash-tier agents today, and Model B provides no evidence that a "3.5 Flash Cyber" model or "3.5 Pro partner testing" actually appear in the source article rather than being hallucinated additions.”

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @google 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/gemini-3-6-flash-red…] indexed:0 read:1min 2026-07-22 ·