New AI University AI Topics
← AI News

Google Cuts AI Costs by 65% with New Gemini 3.6 Flash Model

Source: VentureBeat AI

Summary

  • Google's new Gemini 3.6 Flash model cuts AI agent token costs by up to 65% on long horizon engineering tasks.
  • The model is priced at $1.50 per one million input tokens and $7.50 per one million output tokens through Google's API.
  • This is a significant reduction compared to the previous Gemini 3.5 Flash model, which costs $1.50/$9.00 per 1M tokens.
  • The new model is designed to make AI agents faster, smarter, and cheaper at scale.
  • Google also released two other new models, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber, which are priced at $0.30/$2.50 and $0.30/$2.50 per million tokens in/out, respectively.
  • Google's new Gemini 3.6 Flash model is priced lower than many other AI models on the market.
  • For example, the Xiaomi MiMo-V2.5 Flash model costs $0.10/$0.30 per million tokens in/out, while the DeepSeek deepseek-v4-flash model costs $0.14/$0.28 per million tokens in/out.
  • The new model is also designed to be more efficient and faster than previous models, making it a more attractive option for businesses and organizations that rely on AI agents.
  • Google's Gemini 3.5 Flash-Lite model is priced at $0.30/$2.50 per million tokens in/out, making it a more affordable option for those who need AI agents but do not require the full power of the Gemini 3.6 Flash model.
  • The Gemini 3.5 Flash Cyber model is priced similarly to the Gemini 3.6 Flash model, but it is designed for more complex tasks and may be more suitable for businesses and organizations that require advanced AI capabilities.

Why It Matters

  • Google's new Gemini 3.6 Flash model is a significant development in the field of AI, and it has important implications for businesses and organizations that rely on AI agents.
  • By reducing the cost of AI agent tokens by up to 65%, Google is making it more possible for businesses and organizations to adopt AI agents and use them to automate tasks and improve efficiency.
  • This can lead to increased productivity, improved customer satisfaction, and enhanced competitiveness.
  • The new model is also an important step forward in the development of more efficient and faster AI agents.
  • As AI continues to evolve and become more complex, businesses and organizations will need to have access to more advanced AI capabilities to stay competitive.
  • Google's Gemini 3.6 Flash model is an important step in this direction, and it has the potential to revolutionize the way businesses and organizations use AI agents.

GenAI EXPLAINED

Token Efficiency: Token efficiency refers to the ability of an AI model to process and generate text using the fewest number of tokens (the building blocks of text) possible. The Gemini 3.6 Flash model is designed to be more token-efficient than previous models, which means it can process and generate text using fewer tokens. This makes it faster and more cost-effective.

Token Cost: Token cost refers to the cost of using AI agents to process and generate text. The Gemini 3.6 Flash model is priced at $1.50 per one million input tokens and $7.50 per one million output tokens, which is a significant reduction compared to the previous Gemini 3.5 Flash model.

Long Horizon Engineering Tasks: Long horizon engineering tasks refer to complex tasks that require AI agents to process and generate text over a long period of time. The Gemini 3.6 Flash model is designed to be particularly effective for these types of tasks, which makes it a valuable tool for businesses and organizations that rely on AI agents.