Google Gemini 3.6 Flash launch cuts token use and undercuts 3.5 Flash

Posted by Mahi Gupta
 Google Gemini 3.6 Flash launch cuts token use and undercuts 3.5 Flash

Google’s Gemini 3.6 Flash is the main update here, and the company says it uses 17% fewer output tokens than Gemini 3.5 Flash. That is the part that matters most for teams running AI agent workflows at scale, where every tool call and response can affect API spend and throughput planning.

Gemini 3.5 Flash-Lite hits 350 tokens per second for bulk jobs

Google is also pairing the launch with Gemini 3.5 Flash-Lite, a model it describes as built for high-throughput use rather than heavier reasoning. At 350 tokens per second, it is aimed at low-latency bulk usage in the Gemini API pricing update cycle, where speed can matter more than model depth.

For teams handling large volumes of routine tasks, that kind of speed is usually the bigger draw. It points to a model family that is being split more clearly by workload type, with Flash-Lite handling the faster, lighter end of the range. In practice, that can help developers choose a model based on throughput needs instead of using one general option for everything.

  • Gemini 3.5 Flash-Lite is positioned for high-throughput work.
  • Google says it can reach 350 tokens per second.
  • The focus is low latency bulk usage rather than deeper reasoning.
  • The use case appears tied to efficiency in API-heavy workflows.

CodeMender keeps Gemini 3.5 Flash Cyber on security duty

Gemini 3.5 Flash Cyber is staying inside CodeMender security agent workflows instead of getting a broad public push. Google is limiting access to governments and trusted partners, which suggests the model is being treated as a specialist tool for vulnerability hunting rather than a general-purpose chat model.

That distinction matters because it shows how Google is separating consumer-facing models from more controlled security use cases. Rather than opening the model broadly, the company appears to be keeping it focused on a narrower operational role. For organizations working in security, that usually means tighter access, clearer deployment boundaries, and more specific expectations around what the model is meant to do.

The update also fits with the broader pattern in Google’s Gemini 3.5 line. One model is being pushed for lower token use, another for speed, and another for controlled security work. Taken together, that gives the lineup a more practical shape, with each model aimed at a different kind of workload.

  • Gemini 3.5 Flash Cyber remains tied to CodeMender workflows.
  • Access is limited to governments and trusted partners.
  • The model is being treated as a specialist security tool.
  • Google is not positioning it as a broad public chat model.

For developers and enterprise teams, the key takeaway is simple. Google is making the Gemini 3.5 family more task-specific, and the biggest efficiency gain so far is the 17% reduction in output tokens on Gemini 3.6 Flash. That could matter most where scale, latency, and cost control all sit in the same decision.

Mahi Gupta

Mahi Gupta

author

✉ mahigupta708076@gmail.com

Hi, I'm Mahi Gupta the Tech Writer at JhatpatLo. I write about smartphones, Android, Apple, AI, gadgets, software updates, and consumer technology. My goal is to make technology easy to understand by publishing accurate, well-researched, and reader-friendly content.Through JhatpatLo, I help readers stay updated with the latest tech news, buying guides, comparisons, and practical tips.