GKRootWire
Security ICE Signs $2M Deal for Zero-Click Phone Hacking ToolSecurity Attackers Exploit Critical Elementor Pro Bug to Hijack WordPress SitesAI ChatGPT Goes Down, Serves 404 Errors to UsersAI ChatGPT and Codex Suffer Widespread OutageAI Google DeepMind's WeatherNext 3 Sharpens AI Weather ForecastingAI Google's New AI Weather Model Sharpens Storm ForecastsSecurity ICE Signs $2M Deal for Zero-Click Phone Hacking ToolSecurity Attackers Exploit Critical Elementor Pro Bug to Hijack WordPress SitesAI ChatGPT Goes Down, Serves 404 Errors to UsersAI ChatGPT and Codex Suffer Widespread OutageAI Google DeepMind's WeatherNext 3 Sharpens AI Weather ForecastingAI Google's New AI Weather Model Sharpens Storm Forecasts
AI

Google's New Gemini 3.8 Flash Thinks Harder, But That Thinking Isn't Free

Google's latest Flash model adds extra reasoning steps and iterative tool calls, which could quietly inflate your API bill despite unchanged sticker pricing.

Google has rolled out Gemini 3.8 Flash, replacing 3.7 Flash just weeks after that model's debut. The headline change isn't a price cut or a benchmark flex — it's that the model is designed to "work harder" by running more reasoning steps and calling external tools repeatedly when tackling complex tasks.

Pricing stays the same as 3.7 Flash: $0.75 per million input tokens and $3.75 per million output tokens. But Google itself cautions that the model may burn through more tokens to hit peak performance, particularly at higher "effort" settings, meaning actual costs could climb even though the rate card didn't change. Developers who want predictable, lower token usage can stick with 3.7 Flash instead.

The release has already drawn some early scrutiny from developers testing the tradeoffs in practice.

Why it matters: This is a preview of how AI vendors may increasingly compete: not on transparent per-token pricing, but on opaque 'effort levels' that shift real costs onto usage rather than list price. Teams building on Gemini should budget for token volume, not just rate cards, and treat model upgrades as cost experiments rather than free wins.

Sources: The Verge