Google's lighter Gemini models reveal AI’s new race Google launched Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber on Tuesday, offering lighter models with lower cost and latency for developers building AI agents. Gemini 3.6 Flash uses up to 17% fewer output tokens than its predecessor, while 3.5 Flash Cyber is fine-tuned for cybersecurity vulnerabilities. Google also previewed CodeMender, a security agent that uses 3.5 Flash Cyber to scan and fix code vulnerabilities in internal codebases including Chrome and Android. As the race for token efficiency heats up, Google is making its move: three new models designed to help developers do more for less. On Tuesday, Google launched Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Cyber. The new additions to the Flash lineup are meant to provide developers with lighter models that have lower cost and latency, but don't compromise efficiency. That's a balance that is especially necessary to build, deploy and scale AI agents. Each model has its own strengths. Gemini 3.6 Flash builds on Gemini 3.5 Flash https://www.thedeepview.com/articles/gemini-3-5-and-spark-get-google-in-the-agent-race , announced in May at Google I/O. The company incorporated developer and customer feedback to improve coding, knowledge, and multimodal performance, but all at a lower cost per token. According to Google, on the Artificial Analysis Index, 3.6 Flash uses up to 17% fewer output tokens than its predecessor, underscoring its efficiency. A quick look at the other models: 3.5 Flash-Lite : Google shares that it is its "fastest and most effective 3.5-class model," meant for tasks that require both low latency and high throughput. 3.5 Flash Cyber : Built on top of 3.5 Flash, it was fine-tuned to find and fix cybersecurity vulnerabilities at a lower price per token than its larger counterparts. Notably absent was Google's flagship Gemini 3.5 Pro, which has reportedly been delayed https://www.cnbc.com/2026/07/16/alphabet-stock-gemini-3-5-pro-ai.html and is months behind schedule due to performance issues and not living up to the coding prowess of rivals Anthropic and OpenAI, according to the report. Google did mention in this release that Gemini 3.5 Pro is currently testing with partners and plans to make it broadly available "as soon as it’s ready." Google also launched a preview of CodeMender, its security agent managed and hosted by Google, which can scan code to find vulnerabilities and fix them using the brand new 3.5 Flash Cyber. The company shares that CodeMender is already finding and fixing vulnerabilities in Google's internal codebases, including Chrome, Android, Cloud, Ads and YouTube. Gemini 3.5 Flash Cyber will be available exclusively via CodeMender to governments and trusted partners as part of a limited-access program. Meanwhile, everyone can access Gemini 3.6 Flash and 3.5 Flash Lite in the Gemini app, and 3.5 Flash-Lite is also rolling out to Google Search. Developers can access them in the Gemini API via Google AI Studio and Android Studio. Enterprises can access them in the Gemini Enterprise Agent Platform. Gemini 3.6 Flash is also available in Google Antigravity and the Gemini Enterprise app. Our Deeper View AI agents, although prolific in enterprise environments, require a lot of processing to run in the background. As a result, demand is higher than ever for models that are highly capable while keeping token costs low, especially since most AI enthusiasts run multiple agents at the same time. Google is good at listening to developer requests and catering to the community, and this announcement is an example of that. Lastly, this shows the same trend I have been writing about nearly every day for the past two weeks: an industry committed to decreasing AI costs. The biggest challenge for Google is the continued delay of its flagship Pro model, which would allow its coding tools to match the runaway success of Claude Code and Codex among AI enthusiasts and builders.