Google DeepMind releases Gemini 3.6 Flash and 3.5 Flash Lite cutting AI agent token costs by up to 65 percent
> cd .. / HUB_EDITORIALE
News

Google DeepMind releases Gemini 3.6 Flash and 3.5 Flash Lite cutting AI agent token costs by up to 65 percent

[2026-07-22] Author: Meteora Web Redazione
> share
Zenithby Meteora Web The operating system for your business. Social, clients, bookings and invoices in one platform. Gyms, barbers, professionals. Discover Zenith Free demo · no card

Google DeepMind today released three new proprietary AI models designed to optimize the cost and performance of AI agents. The Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber models represent a significant step toward faster and cheaper agents at scale. Pricing ranges from $0.30 per million input tokens for the Lite model to $7.50 per million output tokens for the main Flash version.

Three new models with competitive pricing and up to 65 percent token savings

Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens via API. Gemini 3.5 Flash-Lite costs only $0.30 per million input tokens and $2.50 per million output tokens, making it one of the cheapest models on the market. Compared to the previous Gemini 3.5 Flash at $1.50 input and $9.00 output, the new models offer significant savings. According to the independent Artificial Analysis Index, Gemini 3.6 Flash reduces output token usage by 17 percent over its predecessor, while on specific benchmarks like DeepSWE for long-horizon software engineering tasks, savings reach 65 percent. This means the model requires fewer reasoning steps and tool calls to complete complex workflows, further lowering effective costs for enterprises. Google stated that the models are designed to use fewer tokens overall, offering superior value over list prices.

Sponsored Protocol

Improved benchmark performance and enhanced safety

Technological improvements translate into higher scores on several tests. Gemini 3.6 Flash achieves 49 percent on DeepSWE, up from 37 percent in version 3.5. On MLE-Bench for machine learning engineering, the score rises to 63.9 percent from 49.7 percent. The model also integrates computer use as a client-side tool via API, with an OSWorld-Verified score of 83 percent. For safety, Google has deployed Frontier Safety protections that harden the model against jailbreaks and mitigate risks in chemical, biological, radiological, and nuclear domains, as well as cyber misuse. The models are trained to minimize refusals for beneficial uses, balancing security and practical utility.

Sponsored Protocol

Three variants for specific needs: coding, speed, and cybersecurity

Google divided the offerings into three distinct products. Gemini 3.6 Flash serves as the workhorse for complex coding, knowledge work, and multimodal processing. It is used for document analysis, chart parsing, and report drafting. Gemini 3.5 Flash-Lite is optimized for high throughput and minimal latency: it processes 350 tokens per second, twice as fast as the previous Gemini 3.1 Flash-Lite. It is ideal for agentic search and massive document processing. Gemini 3.5 Flash Cyber is a specialized cybersecurity model trained to find and fix vulnerabilities. It will be available exclusively to governments and trusted partners via CodeMender, Google's bug-fixing agent. Unlike open-source models, all new Gemini models are closed source and accessible only through official API, creating a dependency on Google's infrastructure.

Awaiting Gemini 3.5 Pro while Gemini 4 pre-training begins

Despite the enthusiasm for the new models, developers noted the absence of the flagship Gemini 3.5 Pro, which Google had anticipated for summer. Logan Kilpatrick of Google responded on X that the model is being tested with partners and will be released when ready. Meanwhile, Google confirmed that pre-training for Gemini 4 has already started. This release signals that the immediate future of AI lies in agentic capabilities—systems that operate autonomously over extended periods. The Flash models represent a step toward more agile and efficient agents, reducing operational costs for businesses. For more insights on AI agent evaluation, refer to the article on LangChain, Conviva, CoreWeave, and for a comparison on cybersecurity, Apple Patches Hide My Email Bug offers relevant context. For further details, see the original article on VentureBeat.

Sponsored Protocol

Source: https://venturebeat.com/technology/googles-gemini-3-6-flash-model-cuts-ai-agent-token-costs-by-up-to-65-on-long-horizon-engineering-tasks-and-3-5-pro-is-on-the-way

> share
Meteora Web Redazione

> AUTHOR_EXTRACTED

Meteora Web Redazione

La redazione di Meteora Web Agency: ingegneri informatici e professionisti del digitale che pubblicano ogni giorno news e approfondimenti su tecnologia, software, marketing e innovazione.
[ Read Full Dossier ]

> METEORA_WEB // DIGITAL AGENCY

We build the digital presence your business deserves.

Websites, social media, online advertising, e-commerce and high-performance hosting, engineered with method by computer engineers in Sciacca, for all of Italy.

> MW_JOURNAL

> READ_ALL()