Google has released a new series of Gemini models, and the strategic direction is clear: the focus has shifted decisively to the Flash series rather than the Pro series. The reason is structural. Pro models are expensive. They require massive computational resources—GPUs, cooling, electricity, processors. Running a powerful AI model for an hour demands a significant budget, and that is unsustainable at scale. The AI race has therefore split into two dimensions: horizontal, pursuing the most powerful model possible, and vertical, pursuing the most powerful model at the lowest possible cost. The Flash series is Google's answer to that vertical challenge. Three new models have been released: Gemini 3.6 Flash, Gemini 3.5 Flash Lite, and Gemini 3.5 Flash Cyber. Each serves a distinct purpose. Here is what they are, how they compare, and how to use them.
The new Flash generation is built around three core improvements. Response latency has been reduced by up to 40% compared to previous models. This means answers arrive significantly faster—not because the model is simpler, but because it processes and understands input more efficiently. Cost efficiency has been dramatically improved, enabling the processing of massive datasets at very low operational cost, making these models viable for large-scale applications. Advanced security specialization has been embedded into one of the three models specifically, trained for cybersecurity analysts.
Gemini 3.6 Flash. This is positioned as the ideal choice for multimodal tasks—images, video, PDFs, Word documents, and more. It combines complex language understanding with near-instant response times. It is the general-purpose workhorse of the new lineup.
Gemini 3.5 Flash Lite. As the name indicates, this is the lightweight option. It is designed for high-volume query systems that require fast, efficient responses. It is the model that will be set as default on the official Gemini interface. Despite being "Lite," it is not weak—it represents Google's pursuit of the difficult equation: strong capability plus low cost plus high speed.
Gemini 3.5 Flash Cyber. This model is trained specifically for cybersecurity analysts. Its functions include vulnerability detection, malicious code analysis, and security auditing. It is not designed for general public use—it serves a specialized professional audience.
The Cyber variant demonstrates exceptional performance in its domain. On the Metis benchmark, it slightly outperforms competing models. Against GPT-5.6, it is very close—within a few percentage points on the SOL benchmark. Against GPT-5.5 Cyber, the gap is similarly narrow. Where the specialization becomes dramatically visible is on benchmarks like the Epic Sleepy Evaluation. The standard Gemini 3.5 Flash scores 36. The Cyber variant scores 72—double the performance. This is the result of targeted training on cybersecurity-specific data. The model excels in its domain precisely because it is not asked to be good at everything.
The Flash Lite model achieves its speed through a technical capability called "53 output tokens per second"—meaning it understands input rapidly, which is why it responds rapidly. Thinking levels are adjustable, allowing users to control the depth of reasoning versus speed of response. The cost is positioned to be accessible compared to other models in the market.
The Flash Lite model has already demonstrated capabilities that go beyond simple question-answering. In demonstrations, it built an agentic puzzle maker—a tool for generating puzzles. It showed the ability to combine images in creative ways. These are not trivial tasks. The model being fast does not mean it is weak. Google's engineering challenge was precisely to avoid that tradeoff: to deliver a model that competes with the current market leaders in capability while costing less and running faster.
All three models are now available across Google's platforms. On the official Gemini interface, 3.5 Flash Lite and 3.6 Flash are activated by default. On Google AI Studio, both 3.6 Flash and 3.5 Flash Lite are available. On Antigravity, the full range is accessible—3.6 Flash and 3.6 Flash Medium are both listed. The models have been officially launched and deployed across Google's entire ecosystem.
The practical implication for business is significant. Faster, cheaper models that maintain competitive capability mean that AI-powered applications, tools, and automations become more accessible to build and more economical to run at scale. Every improvement in this direction directly benefits anyone building products on top of these models.