AB
AiBoss
project

Gemini 3.5 Flash-Lite - A lightweight AI model released by Google.

The Gemini 3.5 Flash-Lite is Google's fastest and cheapest Gemini 3.5 series model, designed for high-throughput, low-latency production-grade agent workflows.

What is Gemini 3.5 Flash-Lite?

Gemini 3.5 Flash-Lite is Google's fastest and cheapest Gemini 3.5 series model, designed for high-throughput, low-latency production-grade agent workflows. The model supports adjustable inference speeds, outputs up to 350 tokens/s, and is priced at only $0.30/$2.50 per million tokens. It surpasses its predecessor, 3 Flash, in several agent and code benchmarks, making it suitable for batch document processing, sub-agent collaboration, and real-time interaction scenarios.

Main functions of Gemini 3.5 Flash-Lite

  • High-speed outputThe fastest in the 3.5 series, with Artificial Analysis achieving a measured output of 350 tokens per second.
  • Adjustable inference settingsSupports low/medium/high thinking modes. For simple tasks, use the lowest setting to ensure speed, and for complex sub-tasks, use the highest setting to ensure quality.
  • Built-in computer operationComputer Use is a built-in tool that supports Agent tasks out of the box.
  • Quality improvementCompared to Flash-Lite 3.1, the code, long context, and real task execution capabilities have been greatly improved.

The technical principle of Gemini 3.5 Flash-Lite

  • Architectural foundationBased on the streamlined and optimized Gemini 3.5 Flash architecture, it retains core capabilities while minimizing inference overhead.
  • Dynamic thinking mode: Switch between latency and quality on demand through configurable thinking levels.
  • High throughput optimizationOptimizes parallel generation efficiency for scenarios such as batch document processing and agent search, achieving a peak output of 350 tokens/s.

How to use Gemini 3.5 Flash-Lite

  • DevelopersIt supports flexible configuration of thinking levels through Google AI Studio and Gemini API calls.
  • Enterprise usersIntegrated into the Gemini Enterprise Agent Platform, it is suitable for high-concurrency production traffic.
  • Regular usersIt has been gradually rolled out to consumer-facing scenarios such as Google Search.

The core advantages of Gemini 3.5 Flash-Lite

  • Speed and cost dual extremesWith a generation speed of 350 tokens/s and a pricing of $0.3/$2.5, the cost of large-scale tasks is extremely low.
  • Winning big with small investmentIt outperformed the older 3 Flash in SWE-Bench Pro and OSWorld-Verified (74.0% vs 65.1%).
  • Flexible thinking gearDevelopers can freely switch the inference depth according to the task complexity, without having to pay unnecessary computational overhead for simple tasks.
  • Agent nativeBuilt-in Computer Use and multiple thinking modes, designed specifically for scenarios such as sub-agent collaboration and batch data processing.

Comparison of Gemini 3.5 Flash-Lite with similar competing products

Comparison Dimensions Gemini 3.5 Flash-Lite Gemini 3.1 Flash-Lite
Terminal-Bench 2.1 54.0% 31.0%
GDPval-AA v2 1140 642
GDM-MRCR v2 72.2% 60.1%
Output speed 350 token/s Not disclosed
Enter price $0.30/million tokens higher
Output Price $2.50/million tokens higher

Application scenarios of Gemini 3.5 Flash-Lite

  • High-throughput Agent SearchLarge-scale web page retrieval, information aggregation and summary generation, with low-latency response to user queries.
  • Batch document processing: Massive data pipelines such as receipt translation, e-commerce product feature extraction, and batch contract parsing.
  • Sub-Agent CollaborationAs a secondary agent to the primary agent (such as Flash 3.6), it can quickly generate multiple versions of solutions (such as 25 web design concepts) for selection.
  • Real-time interactive applicationsScenarios requiring rapid response, such as customer service robots, real-time recommendation systems, and low-latency dialogue interfaces.
  • Lightweight game/creative generationRapidly iterate to generate mini-game prototypes and visual assets, and optimize them in real time with human feedback.