Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
Google Blog 2026-07-21 15:00:00
Context: Google has introduced new AI models under its Gemini series, including Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, designed to enhance efficiency, latency, and reliability for building AI agents at scale. These models aim to improve token efficiency, reduce latency, and offer better price-to-performance ratios for developers and customers. The new models are built to support various applications, including coding, knowledge work, and cybersecurity.
Key Facts
- Gemini 3.6 Flash consumes 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index and offers improved token efficiency, making it more cost-effective with a price of $1.50/1M input tokens and $7.50/1M output tokens.
- Gemini 3.5 Flash-Lite is the fastest model in the 3.5 series, running at 350 output tokens/s, priced at $0.3/1M input tokens and $2.5/1M output tokens, and offers significantly better quality than 3.1 Flash-Lite.
- Gemini 3.6 Flash demonstrates performance gains compared to 3.5 Flash across various use cases, including parsing financial data, executing code migrations, and developing a photographic texture extractor for 3D workflows.
- Gemini 3.5 Flash Cyber is built on top of 3.5 Flash and fine-tuned for finding and fixing cybersecurity vulnerabilities at a lower price per token than larger models, and will be exclusively available to governments and trusted partners via CodeMender as part of a limited-access pilot program.
- Gemini 3.5 Flash-Lite outperforms 3 Flash on various evaluations, including SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%), making it a faster and more capable option for workloads on both 2.5 and 3 Flash.