Back to articles
Model Releases

Google’s Gemini 3.8 Flash Adds More Reasoning—And May Use More Tokens

3 min read

Google has accelerated the release cycle for its Gemini Flash family with Gemini 3.8 Flash, arriving only a few weeks after its predecessor. The new model’s main pitch is not simply lower latency or a cheaper API. Instead, Google says it is built to “work harder” on difficult problems by taking additional reasoning steps and calling tools iteratively when a task requires more than one action.

Key points

  • The listed price is unchanged, but the bill may not be: Gemini 3.8 Flash launches at $0.75 per million input tokens and $3.75 per million output tokens, matching the introductory pricing of 3.7 Flash. Google warns that the model may consume more tokens to maximize performance, particularly at higher effort levels.
  • The focus is on complex, agentic work: Google highlights improvements in software engineering and autonomous AI agents. The model reportedly outperforms its predecessor and other frontier systems on DeepSWE v1.1, Vals Finance Agent V2, and Harvey’s Legal Agent benchmark.
  • Task-level economics matter more than token rates: Early analysis from Artificial Analysis described the model as highly competitive for its intelligence level, while also noting that output tokens per task and the number of turns in agent evaluations have increased. This means a stable API price does not necessarily translate into a stable cost per job.
  • Safety is part of the release story: Google says the model includes safeguards aimed at misuse involving chemical, biological, radiological, and nuclear risks, as well as cyber offense. Gemini 3.8 Flash Cyber is being made available to governments and trusted partners through the Fairwind Program.

Why the update matters

Flash models have traditionally represented the faster and more economical side of a model lineup. Gemini 3.8 Flash suggests that Google is broadening that role. It keeps a relatively low per-token price while allowing the system to spend more computation and perform more interactions when a problem is difficult. For coding, research analysis, and document-heavy professional workflows, that trade-off could produce more reliable results.

For developers, however, the important metrics are no longer limited to input and output rates. Token totals, tool-call counts, latency, and the number of agent turns may determine the real cost of an application. Teams that need predictable spending can continue using Gemini 3.7 Flash to reduce token consumption. Those willing to pay for additional reasoning may find 3.8 Flash better suited to agentic applications.

The Cyber release also shows Google pairing stronger capabilities with tighter access controls in sensitive environments. The model’s long-term competitiveness will therefore depend on more than benchmark scores. It will need to balance capability, operating cost, speed, and safeguards in real deployments.

Source: The Verge AI

Comments

Checking sign-in status...

Loading comments...

Related articles