project
Gemini 3.5 Flash-Lite - A lightweight AI model released by Google.
The Gemini 3.5 Flash-Lite is Google's fastest and cheapest Gemini 3.5 series model, designed for high-throughput, low-latency production-grade agent workflows.
What is Gemini 3.5 Flash-Lite?
Gemini 3.5 Flash-Lite is Google's fastest and cheapest Gemini 3.5 series model, designed for high-throughput, low-latency production-grade agent workflows. The model supports adjustable inference speeds, outputs up to 350 tokens/s, and is priced at only $0.30/$2.50 per million tokens. It surpasses its predecessor, 3 Flash, in several agent and code benchmarks, making it suitable for batch document processing, sub-agent collaboration, and real-time interaction scenarios.
Main functions of Gemini 3.5 Flash-Lite
-
High-speed outputThe fastest in the 3.5 series, with Artificial Analysis achieving a measured output of 350 tokens per second.
-
Adjustable inference settingsSupports low/medium/high thinking modes. For simple tasks, use the lowest setting to ensure speed, and for complex sub-tasks, use the highest setting to ensure quality.
-
Built-in computer operationComputer Use is a built-in tool that supports Agent tasks out of the box.
-
Quality improvementCompared to Flash-Lite 3.1, the code, long context, and real task execution capabilities have been greatly improved.
The technical principle of Gemini 3.5 Flash-Lite
-
Architectural foundationBased on the streamlined and optimized Gemini 3.5 Flash architecture, it retains core capabilities while minimizing inference overhead.
-
Dynamic thinking mode: Switch between latency and quality on demand through configurable thinking levels.
-
High throughput optimizationOptimizes parallel generation efficiency for scenarios such as batch document processing and agent search, achieving a peak output of 350 tokens/s.
How to use Gemini 3.5 Flash-Lite
-
DevelopersIt supports flexible configuration of thinking levels through Google AI Studio and Gemini API calls.
-
Enterprise usersIntegrated into the Gemini Enterprise Agent Platform, it is suitable for high-concurrency production traffic.
-
Regular usersIt has been gradually rolled out to consumer-facing scenarios such as Google Search.
The core advantages of Gemini 3.5 Flash-Lite
-
Speed and cost dual extremesWith a generation speed of 350 tokens/s and a pricing of $0.3/$2.5, the cost of large-scale tasks is extremely low.
-
Winning big with small investmentIt outperformed the older 3 Flash in SWE-Bench Pro and OSWorld-Verified (74.0% vs 65.1%).
-
Flexible thinking gearDevelopers can freely switch the inference depth according to the task complexity, without having to pay unnecessary computational overhead for simple tasks.
-
Agent nativeBuilt-in Computer Use and multiple thinking modes, designed specifically for scenarios such as sub-agent collaboration and batch data processing.
Comparison of Gemini 3.5 Flash-Lite with similar competing products
| Comparison Dimensions | Gemini 3.5 Flash-Lite | Gemini 3.1 Flash-Lite |
|---|---|---|
| Terminal-Bench 2.1 | 54.0% | 31.0% |
| GDPval-AA v2 | 1140 | 642 |
| GDM-MRCR v2 | 72.2% | 60.1% |
| Output speed | 350 token/s | Not disclosed |
| Enter price | $0.30/million tokens | higher |
| Output Price | $2.50/million tokens | higher |
Application scenarios of Gemini 3.5 Flash-Lite
-
High-throughput Agent SearchLarge-scale web page retrieval, information aggregation and summary generation, with low-latency response to user queries.
-
Batch document processing: Massive data pipelines such as receipt translation, e-commerce product feature extraction, and batch contract parsing.
-
Sub-Agent CollaborationAs a secondary agent to the primary agent (such as Flash 3.6), it can quickly generate multiple versions of solutions (such as 25 web design concepts) for selection.
-
Real-time interactive applicationsScenarios requiring rapid response, such as customer service robots, real-time recommendation systems, and low-latency dialogue interfaces.
-
Lightweight game/creative generationRapidly iterate to generate mini-game prototypes and visual assets, and optimize them in real time with human feedback.