Google announced three fresh Gemini models on the heels of its May I/O showcase, where Gemini 3.5 Flash stole the spotlight. The company says the new lineup supersedes the older model, which it has already deprecated, and adds two specialized variants aimed at efficiency and security.
Gemini 3.6 Flash is the flagship replacement. Google cites a jump from 37% to 49% on the DeepSWE coding benchmark and a modest rise from 78.4% to 83% on the OSWorld computer‑use test. Those gains come alongside a 17% reduction in token usage, a metric that directly translates into lower operating costs for developers.
Pricing reflects the efficiency push. The API now costs $1.50 per million input tokens and $7.50 per million output tokens, down from $9 for output on the previous 3.5 Flash. Google argues the cheaper, faster model will let developers build more complex agentic workflows without inflating budgets.
Alongside the flagship, Google released Gemini 3.5 Flash Lite, billed as the most efficient modern AI model in the Gemini family. The Lite version can handle about 350 tokens per second, a speed the company says rivals frontier models from a year ago. Its price tag sits at $0.30 per million input tokens and $2.50 per million output tokens, slightly higher than the earlier 3.1 Flash Lite but still positioned as a cost‑effective option for large‑scale deployments.
The third addition, Gemini 3.5 Flash Cyber, marks Google’s first Gemini model built with cybersecurity in mind. While Google has not disclosed benchmark results for the new security‑focused variant, the launch signals an intent to embed AI capabilities into threat detection and response workflows.
Notably, the anticipated Gemini 3.5 Pro, originally slated for a June release, remains on the back burner. Google did not provide a new timeline, leaving developers who were counting on the Pro version to adjust their roadmaps.
Overall, Google frames the updates as a response to user feedback on Gemini 3.5 Flash, which fell short of expectations in code generation and cost efficiency. By tightening token consumption, sharpening performance scores and diversifying the model family, the company hopes to keep its AI offering competitive as enterprises grapple with rising AI token expenses.
Questo articolo è stato scritto con l'assistenza dell'IA.
News Factory APP - notizie agentiche per potenziare il tuo SEO e AEO.