Google turns Flash models toward cost

Google's July 21 launch of three Flash models looks like another Gemini update. Read the pricing, speed and intended uses together, and the move is clearer: Google is trying to make agentic and enterprise AI cheaper, faster and more specific to the work it performs.

The lineup separates three jobs. Gemini 3.6 Flash is the broader workhorse, Gemini 3.5 Flash-Lite is tuned for low cost and high throughput, and Gemini 3.5 Flash Cyber is aimed at security work. That split says model competition is moving from a single leaderboard race into workload design.

Cheap models carry frequent calls

Google says the new Flash-Lite has 17% faster real-time output than the previous generation and reaches roughly 350 tokens per second in independent testing. DeepMind's model page lists prices of $0.30 per million input tokens and $2.50 per million output tokens.

For ordinary chat, users notice answer quality. For agents, customer routing, browser automation and bulk document workflows, the question is whether hundreds or thousands of small calls can stay inside latency and budget limits.

Gemini 3.6 Flash takes a different role. Google says it improves coding, agentic tasks and multimodal understanding over 3.5 Flash, with listed pricing of $1.50 per million input tokens and $7.50 per million output tokens. It is closer to a daily workhorse, while Flash-Lite carries lower-risk high-frequency steps.

Security becomes a separate model

Flash Cyber is the sharper part of the launch. Google positions it for security analysis, vulnerability research, malware analysis and incident response rather than general chat.

Security work has its own metrics. A model has to read logs, code and alerts while controlling false positives and missed findings. Google says Flash Cyber posts gains of up to 65% over 3.5 Flash on selected cybersecurity benchmarks, and DeepMind lists results across CTI-MCQ, CTI-RCM, MARS-EVAL and CyberSecEval.

Google CEO Sundar Pichai has described the AI opportunity as large as it gets.

In this launch, that opportunity is not only a larger flagship model. It is also a menu of models that lets teams choose the right cost, latency and risk profile for each part of a workflow.

Customers are buying a balance

Google's model page includes a useful line from Figma engineering leader Vincent van der Meulen: when Figma evaluates models for Figma Make, it looks for a balance of quality, speed and cost.

That is the commercial logic behind Flash. Design tools, code assistants, spreadsheet analysis and support systems rarely need the same model for every step. They need enough quality at the right speed and price, with a fallback path when the task becomes more sensitive.

For developers in China and elsewhere, the practical lesson is simple. An agent pipeline often contains search, reading, planning, tool calls, rewriting and checking. Running every step on a flagship model can make demos look good, then make production bills and latency hard to defend.

Real workloads will decide

The next checks are concrete. Flash-Lite's 350 tokens per second figure is a benchmark context, so production latency will depend on prompt size, tool use, network conditions and rate limits. Flash Cyber's 65% improvement is tied to listed security benchmarks, so enterprise use still has to measure false positives, misses and review cost.

Google did not only launch a stronger model. It split different workloads into different products. The next phase of model competition will be fought through bills, latency, security boundaries and workflow fit.

Sources: Google official blog, Google DeepMind model pages, Axios, Business Insider, Artificial Analysis, CocoLoop; checked the three Flash model roles, tokens-per-second scope, per-million-token pricing, cybersecurity benchmark gains, customer quote and availability.