Google launches Gemini 3.6 Flash and 3.5 Flash-Lite, teases Gemini 4
9to5 Google 2026-07-21 15:00:01
Context: Google has announced the launch of Gemini 3.6 Flash and 3.5 Flash-Lite, two new AI models designed to improve performance, efficiency, and coding capabilities. These models follow the last release at I/O 2026 and take into account developer and customer feedback since May. The new models aim to enhance token efficiency, reduce costs, and improve production-ready code generation.
Key Facts
- Gemini 3.6 Flash consumes 17% fewer output tokens compared to its predecessor, while taking fewer reasoning steps and tool calls to accomplish multi-step workflows, and is priced lower at $1.50/1M input tokens and $7.50/1M output tokens.
- Gemini 3.6 Flash delivers higher precision with fewer unwanted code edits and reduced execution loops, generating higher quality and more reliable production-ready code as seen in DeepSWE (49% vs. 37%) and showing significant improvement in ML Research, as seen in MLE Bench (63.9% vs. 49.7%).
- Gemini 3.5 Flash-Lite is designed for high-throughput and low-latency tasks, offering significantly better quality than 3.1 Flash-Lite, with pricing at $0.30/1M input tokens and $2.50/1M output tokens, and outperforming 3 Flash in SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%).
- The knowledge cutoff date for Gemini 3.6 Flash advances from January 2025 to March 2026, and its computer use capabilities improve from 78.4% on OSWorld-Verified to 83%, with a score of 1421 on GDPval-AA.
- Gemini 3.5 Flash Cyber, built on the foundation of Gemini 3.5 Flash, is designed for finding and fixing security vulnerabilities, with access initially available for governments and trusted partners as part of a limited-access pilot program.