Google DeepMind has unveiled three new proprietary AI models that promise to significantly lower operational costs for businesses. Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber are among the most token-efficient models ever created by the Mountain View giant. The goal is to make AI agents faster, smarter, and cheaper at scale.
Gemini 3.6 Flash is priced at $1.50 per million input tokens and $7.50 per million output tokens via API. Gemini 3.5 Flash-Lite costs just $0.30 for input and $2.50 for output, positioning it as an ultra-low-cost option for high-volume workloads. Compared to the previous Gemini 3.5 Flash ($1.50/$9.00) and Gemini 3.1 Pro Preview ($2.00/$12.00), the savings are substantial. However, the entry-level Gemini 3.1 Flash-Lite remains the cheapest at $0.25/$1.50, though it is twice as slow as the new Flash-Lite.
Sponsored Protocol
Record token efficiency with up to 65 percent reduction on software benchmarks
According to the Artificial Analysis Index, Gemini 3.6 Flash reduces output token usage by 17 percent versus its predecessor. On specialized benchmarks like DeepSWE, which measures multi-step engineering tasks, savings reach 65 percent. This means the model requires fewer reasoning steps and tool calls to achieve the same result. Google has optimized internal verbosity, making the model more direct and less wasteful. Both Flash models offer a 1-million-token context window and a 64,000-token output limit, with a knowledge cutoff of March 2026.
Improved performance on coding, knowledge work, and computer use
Gemini 3.6 Flash shows clear progress: on DeepSWE it jumps from 37 percent to 49 percent, on MLE-Bench from 49.7 percent to 63.9 percent, and on OSWorld-Verified (computer use) from 78.4 percent to 83.0 percent. The GDPval-AA v2 benchmark also rises from 1349 to 1421. Google has integrated computer use directly into the Gemini API, enhancing autonomous agents.
Sponsored Protocol
Three models for three distinct use cases: coding, speed, and cybersecurity
Gemini 3.6 Flash is the workhorse for complex programming, document analysis, and reports. Gemini 3.5 Flash-Lite targets environments where minimal latency is critical: it processes 350 tokens per second, twice as fast as the previous model, and is ideal for agentic search and massive document processing. Gemini 3.5 Flash Cyber is a specialized model for cybersecurity, trained to find and fix vulnerabilities. It will be available exclusively to governments and trusted partners via CodeMender, Google's bug-fixing agent.
Immediate availability and absence of the awaited Pro model
Gemini 3.6 Flash and 3.5 Flash-Lite are already available via Gemini API, Google AI Studio, Android Studio, the consumer Gemini app, and Google Search. As usual, these are closed-source models accessible only through official API. Developers have noted the absence of the more powerful Gemini 3.5 Pro, which Google had hinted for summer. Logan Kilpatrick of Google stated on X that the model is being tested with partners and will be released as soon as ready. In the meantime, Google's strategy focuses on efficient AI agents, reducing operational costs even in cybersecurity scenarios. For a broader comparison, see the Wikipedia page on Gemini.
Sponsored Protocol