Skip to content
research

Google's New AI Is Quietly Winning

The AI race seems dominated by OpenAI and Anthropic, who trade blows for the top benchmark spots. But Google's latest model beats them on key engineering tasks for a fraction of the cost, changing the calculus for developers and businesses.

Aki Tanaka
Google's New AI Is Quietly Winning

Google's Rocky Road to Redemption

Google initially found itself flat-footed when chatGPT burst onto the scene, forcing the company to drop everything and scramble to catch up. About a year and a half ago, Gemini 2.5 pro emerged as a brief beacon, hailed as "the best model on the planet" and the first to simulate a full Rubik's's Cube. This leadership proved fleeting, however, as openAI and Anthropic quickly pulled ahead, regaining their lead.

Subsequent Gemini releases over the past six months, despite posting solid benchmarks, often felt underwhelming in practical application. Users reported a brittleness in real-world usage, with models frequently falling short of expectations and leaving much to be desired. This sentiment reflected Google's struggle to consistently translate raw computational power into reliable, real-world performance.

Suddenly

Winning Where It Counts Most

Google’s latest iteration, Gemini 3.8 Flash, demonstrates a targeted prowess that sets it apart. The model achieved a striking 73.7% on the Deep SWE v1.1 benchmark, a critical measure for long-horizon software engineering tasks. This performance effectively ties Claude Opus 5 and notably surpasses GPT 5.6 Sol, positioning Gemini 3.8 Flash as a top contender in complex coding capabilities.

This specialist strength, however, contrasts sharply with its general knowledge performance. On the GDP Val benchmark, which assesses real-world knowledge work like PDF extraction and data analysis, Gemini 3.8 Flash scored 1545. This lags significantly behind leading models, with Claude Opus 5 reaching 1824 and GPT 5.6 Sol achieving 1710, indicating it is "just okay" for broad knowledge applications.

Despite its generalist limitations, Gemini 3.8 Flash solidifies its profile as a potent specialist tool by dominating other niche areas. It secured the number one ranking on the Harvey Legal Benchmark with 61.4%, making it the planet's best model for legal work. Furthermore, it led on Terminal Bench 2.1, an agentic terminal coding assessment, with an 89.4% score.

The Unbeatable Price-Performance Ratio

Gemini 3.8 Flash’s introductory pricing represents a profound disruption. Input tokens cost just $0.75 per million, with output tokens at $3.75 per million. This is a mere fraction of competitor rates; Claude Opus 5, for instance, charges $5 and $25 per million respectively, while even the more budget-oriented GPT 5.6 Terra comes in at $2 and $12 per million. Flash, at its initial rate, is 20-30% the price of Terra.

Raw token prices, however, only reveal part of the story. The true measure of value lies in cost per task, which factors in both token price and the model’s operational efficiency—how many tokens it requires to successfully complete a given operation. Gemini 3.8 Flash’s combination of aggressively low token costs and demonstrated performance creates a powerful economic advantage, making sophisticated AI capabilities genuinely accessible for developers and enterprises.

Google notes these are introductory rates, with a planned increase to $1.50 per million input tokens and $7.50 for output. Even at these standard rates, Gemini 3.8 Flash remains exceptionally cost-effective, positioning it as a leading option in its performance class. This pricing strategy democratizes access to high-performing software engineering assistance, as evidenced by its strong Deep SWE scores. For further details on the model's capabilities, explore Introducing Gemini 3.8 Flash and 3.8 Flash Cyber - Google Blog.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

Beyond Benchmarks: Is It Actually Good?

Initial hands-on evaluations of Gemini 3.8 Flash reveal a nuanced performance profile, reflecting its varied benchmark scores. While it demonstrates surprising aptitude in complex generative tasks, such as rendering an intricate 3D Mount Everest map, its output can fall short on simpler creative briefs like website design. This mirrors its top-tier scores on Deep SWE (73.7%) but "just okay" results on GDP Val for general knowledge work.

For businesses, this translates into a critical shift: the era of identifying one universally "best" AI model concludes. Organizations must now rigorously test models like Gemini 3.8 Flash against their unique internal workloads, developing custom benchmarks to discern optimal fit. A model excelling in software engineering or legal tasks may falter in marketing content generation or broader knowledge work.

Gemini 3.8 Flash emerges not as a blanket "GPT-killer," but rather a powerful, economical weapon for targeted applications. Its aggressive introductory pricing, a fraction of competitors like Claude Opus 5 and even GPT 5.6 Terra, coupled with top-tier performance in specific domains, signifies a smart strategic pivot for Google. This model redefines value, focusing on specialized, cost-effective excellence in the dynamic AI race.

Frequently Asked Questions

What is Google's Gemini 3.8 Flash?

Gemini 3.8 Flash is a new, highly efficient AI model from Google designed to offer strong performance on specific tasks, particularly software engineering and legal analysis, at a very low price point.

How does Gemini 3.8 Flash compare to competitors like GPT-5.6 Sol?

On software engineering benchmarks like Deep SWE, Gemini 3.8 Flash performs on par with or slightly better than top models like Claude Opus 5 and GPT 5.6 Sol. However, it scores lower on general knowledge and reasoning tasks.

What are the best use cases for Gemini 3.8 Flash?

Based on benchmark data, it excels at long-horizon software engineering, agentic coding tasks (Terminal Bench), and legal work, where it scored #1 on the Harvey Legal Benchmark. Its low cost also makes it ideal for high-volume, programmatic tasks.

How much does Gemini 3.8 Flash cost?

It has an introductory price of $0.75 per million input tokens and $3.75 per million output tokens. The standard price is planned to be $1.50 (input) and $7.50 (output), which is still significantly cheaper than comparable models from competitors.

Found this useful? Share it.

For builders

Want Stork to write one of these about your product?

Send us a URL. We use the product, form a view, and publish what we actually think — in 8 languages, labeled Sponsored, with no copy approval on your side. That last part is what makes it worth quoting.

See how it works$500 · AI tools & software only