# Google readies Gemini 3.8 Flash as its engineers reportedly prefer it to Claude Opus

> Source: <https://runtimewire.com/article/google-gemini-3-8-flash-coding-claude-kavukcuoglu>
> Published: 2026-09-02 01:58:32+00:00

# Google readies Gemini 3.8 Flash as its engineers reportedly prefer it to Claude Opus

**The reported coding model could arrive 20 days after Gemini 3.7 Flash, though Google has yet to publish benchmarks, pricing or a model card.**

By [RuntimeWire Staff](/author/runtimewire-staff)
· Published

Primary source: [The Wall Street Journal](https://www.wsj.com/tech/ai/new-google-ai-model-said-to-narrow-gap-on-coding-ability-264c6052)

## Why it matters

If independent tests support Google's internal preference, Gemini 3.8 Flash could give developers a lower-cost coding model for high-volume agent work while testing whether DeepMind can sustain its faster release rhythm.

[Google](https://google.com/?ref=runtimewire) could release Gemini 3.8 Flash as soon as September 2, according to [The Wall Street Journal](https://www.wsj.com/tech/ai/new-google-ai-model-said-to-narrow-gap-on-coding-ability-264c6052?mod=rss_Technology&ref=runtimewire). Employees told the Journal that the model, known internally as "Skimaki," carries significantly improved coding abilities. Some said Google engineers preferred it to an Anthropic Opus model during head-to-head tests inside Jetski, Google's internal coding tool.

That preference is the sharpest claim in the report and the least documented. Google has not published the test prompts, results, evaluator methodology or the exact Opus version used in the comparison. Its official [model-card index](https://deepmind.google/models/model-cards/?ref=runtimewire) still listed [Gemini 3.7 Flash](/models/google/gemini-3.7-flash) as the newest numbered Flash model as of September 2, leaving Gemini 3.8's public availability, specifications, pricing and safety evaluations unconfirmed.

The timing carries weight. Google released Gemini 3.7 Flash on August 13, 20 days before the earliest reported date for 3.8. RuntimeWire [reported at the time](/article/google-gemini-3-7-flash-three-week-release-price-cut) that 3.7 had replaced Gemini 3.6 after another three-week run. A September 2 release would give Google three successive Flash generations in about six weeks.

### A speed mandate at DeepMind

[Koray Kavukcuoglu](https://blog.google/authors/koray-kavukcuoglu/?ref=runtimewire), a longtime DeepMind researcher, took operational control of Google DeepMind after an August 5 leadership overhaul moved co-founder Demis Hassabis into the roles of DeepMind chairman and Alphabet chief scientist, as [The Washington Post reported](https://www.washingtonpost.com/technology/2026/08/05/google-top-ai-leader-demis-hassabis-steps-aside-major-shakeup/?ref=runtimewire). Kavukcuoglu now oversees Google's generative AI models and their integration across its products, according to his official biography, while reporting into [Google and Alphabet CEO Sundar Pichai](https://blog.google/authors/sundar-pichai/?ref=runtimewire).

The appointment put an early DeepMind technical leader in charge of turning research into products at a moment when coding agents have become a proving ground for frontier models. Kavukcuoglu joined DeepMind during its early years, founded its deep-learning team and led work connected to Deep Q-Networks and WaveNet. His background is closer to the model-development machinery than to the conventional corporate operating track.

Gemini 3.8 would offer a visible measure of that operating mandate less than a month into the new structure. The rapid sequence suggests Google is using the Flash line as a vehicle for shipping algorithmic and post-training improvements as soon as they clear internal thresholds, rather than saving each gain for a larger flagship release.

Google described [Gemini 3.7 Flash](https://blog.google/innovation-and-ai/models-and-research/gemini-models/introducing-gemini-3-7-flash/?ref=runtimewire) as a coding and agent model shaped by developer feedback and algorithmic improvements. It was itself released three weeks after Gemini 3.6. Google priced 3.7 at an introductory $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026.

The [3.7 model card](https://deepmind.google/models/model-cards/gemini-3-7-flash/?ref=runtimewire) shows why another coding update could arrive so quickly. Gemini 3.7 scored 43.6% on Google's published FrontierCode 1.1 evaluation, ahead of the 42.7% figure listed for [Claude Sonnet 5](/models/azure/claude-sonnet-5). It still trailed OpenAI's [GPT-5.6 Terra](/models/openai/gpt-5.6-terra) on Google's DeepSWE and Terminal-bench evaluations. The table gave Google a credible coding model without establishing a consistent lead across software-engineering tasks.

### The test Google has not shown

An internal preference test inside Jetski measures something useful: whether engineers choose a model while doing real work. It does not establish that Gemini 3.8 writes more correct, secure or maintainable software across different repositories and development environments.

The reported comparison also leaves Google holding the evaluator, tool and test population. Engineers familiar with Gemini's behavior may prefer its speed, interface, instruction style or integration with Google's systems. Any of those could make 3.8 more productive inside Google without proving a general capability advantage over Anthropic.

A public release should make the comparison easier to interrogate. Pricing and latency matter heavily for a Flash model, which Google positions as a workhorse for high-volume developer and enterprise tasks. A modest capability gain at Flash economics could prove more useful to teams running thousands of agentic tasks than a narrow benchmark victory by a slower, more expensive model.

Independent evaluation will also need to look beyond task completion. A [2025 research paper on code-generation benchmarks](https://arxiv.org/abs/2508.13757?ref=runtimewire) argued for measuring qualities including efficiency, maintainability and security alongside functional correctness. Those dimensions become more important as coding agents move from suggesting snippets to editing repositories and running commands with limited supervision.

### Coding agents are the distribution fight

Anthropic and OpenAI have already built coding products around longer-running, supervised agent work. [Claude Code](https://www.anthropic.com/news/enabling-claude-code-to-work-more-autonomously?subjects=announcements&ref=runtimewire) added checkpoints, subagents, hooks and background tasks so developers can delegate larger jobs while retaining a path to inspect or reverse changes. [OpenAI's Codex app](https://openai.com/index/introducing-the-codex-app/?ref=runtimewire) was designed around managing multiple agents working in parallel across projects.

Google has the pieces to distribute a comparable system through its [Gemini API](https://ai.google.dev/gemini-api/docs/latest-model?ref=runtimewire), [Google AI Studio](https://ai.dev/prompts/new_chat?model=gemini-3.7-flash&ref=runtimewire), [Android Studio](https://developer.android.com/studio?ref=runtimewire) and [Google Antigravity](https://antigravity.google/?ref=runtimewire). A stronger Flash model could improve the economics of those products while giving Google more opportunities to collect feedback from developers using Gemini for actual software work.

Kavukcuoglu's immediate challenge is execution across that full chain: model quality, release speed, developer tooling and price. Gemini 3.8's reported internal reception gives him an encouraging starting point. The public model, once documented, will determine whether Google's engineers spotted a broader coding advantage or simply found a model that works particularly well inside Google.
