Google's Gemini 2.5 Model Technology White Paper (PDF file)
The "Gemini 2.5 Model Technology White Paper" introduces Google DeepMind's Gemini 2.5 model family, a new generation of multimodal AI model family, including versions such as Gemini 2.5 Pro and Gemini 2.5 Flash. Gemini...
《Gemini 2.5 The "Model Technology White Paper" introduces Google DeepMind's...Gemini 2.5 model family, the new generationMultimodalAIModel family, includingGemini 2.5 Pro andGemini Versions such as 2.5 Flash.Gemini 2.5 Pro in inference, coding, andMultimodalIt excels in comprehension, supporting context processing of up to 1 million tokens and capable of analyzing 3 hours of video content. This series of models possesses advanced "thinking" capabilities, supporting dynamic allocation of computing resources for high answer accuracy.Gemini 2.5 Flash delivers high-performance inference with low latency and low cost. The model also shows significant improvements in security, multilingual support, and factuality, and is widely used in code generation, educational tools, and creative design, demonstrating…powerfulIts agency capabilities and practical application potential.
Get Google'sGemini 2.5 Model Technology White PaperOriginal PDF file, scan the QR code to follow and reply: 20250621
Introduction
introduce Gemini 2.5 series model family (including) Gemini 2.5 Pro and Gemini 2.5 Flash), emphasizingMultimodalFeatures include long context, reasoning ability, and tool usage.
- Gemini 2.5 ProThe most currentpowerfulThe model supports 1 million token contexts and can handle 3 hours of video content.
- Gemini 2.5 Flash:High efficiencyThe inference model delivers excellent inference capabilities with low computational and latency requirements.
- Gemini 2.5 Flash-LiteGoogle launchedHigh efficiencyLightweightAIThe model supports long contexts with up to 1 million tokens and tool calls, providing high-performance inference with ultra-low latency and low cost, making it suitable for large-scale application scenarios.
Model Architecture, Training, and Dataset
- Model ArchitectureA Transformer architecture based on Sparse Hybrid Experts (MoE) that supports...MultimodalInput (text, image, audio, video). Based on dynamic routing of input tags to subset parameters (experts), the overall model capacity is decoupled from computation and per-tag service cost.Gemini 2.5 Significant progress has been made in large-scale training stability, signal propagation, and optimization dynamics, directly improving performance after pre-training.
- DatasetPre-trained datasets are large-scale, diverse collections of data covering multiple domains and modalities, including publicly available web documents, code, images, audio, and video.
- Training InfrastructureTraining is based on the TPUv5p architecture. Parallel training is performed using synchronous data, distributed across multiple Google TPUv5p accelerators with 8960 chips. Key software training infrastructure improvements include resilient slicing and split-phase SDC detection, enhancing training resilience and efficiency.
- Post-trainingImprovements to supervised fine-tuning (SFT), reward modeling (RM), and reinforcement learning (RL) enhance model performance.
- Thinking abilityThe model improves answer accuracy based on additional inference time (“Thinking”) and supports dynamic allocation of computing resources.
- Capability-specific improvements:
- CodeSignificant improvement in code generation and comprehension capabilities (e.g., LiveCodeBench score increased from 30.5% to 69.0%).
- FactualIntegrate Google search tools to improveMultimodalThe accuracy of the facts.
- Long ContextOptimize contextual retrieval and inference for millions of tokens.
- MultilingualismSupports 400+ languages, with optimized performance for Chinese, Hindi, and other languages.
- Audio and VideoAdded audio generation and video understanding capabilities (such as 3-hour video analysis).
- Agency capabilities (Agent(ic Use Cases):Gemini Deep Research agents perform exceptionally well in complex tasks.
Quantitative Evaluation
- Methodology:CompareGemini 2.5 Model andGemini The performance of model 1.5, andGemini Performance of 2.5 Pro compared to other large language models.
- Core Capability Results:Gemini 2.5 Pro in code (Aider Polyglot 82.2%), mathematics (AILeading in ME 2025 (88.0%) and long-context tasks (LOFT 87.0%). Compared to others.Large Model(like GPT-4o、Claude 4)Gemini 2.5 ProMultimodalIt performs best in factual tasks.
- Audio/Video EvaluationIt achieves state-of-the-art (SOTA) performance in benchmarks such as FLEURS (speech recognition) and VideoMME (video understanding).
Application Cases
- Gemini Plays PokémonIndependent developers use Gemini The 2.5 Pro version completes Pokémon Blue, demonstrating its ability to plan long-term tasks and perform complex reasoning.
- Other capabilities demonstrated (What Else Can) Gemini 2.5 Do?Transforming scripts into interactive tools, generating SVG from images, creating educational applications, etc.
- Google product integration (Gemini (in Google Products)Applied to Google Search (AI Products such as Overviews and NotebookLM (podcast generator).
Safety, Security, and Responsibility
- Our ProcessSecurity assessment framework, includingautomaticRed team testing and external expert review.
- Policies and DesiderataWe prohibit the generation of harmful content (such as violence and medical errors) and strive for responsiveness, helpfulness, and neutrality.
- Safety TrainingSecurity is optimized through data filtering and reinforcement learning.
- Evaluation Results:Gemini 2.5 Reduced policy violations by 8.2% compared to the previous model, while improving responsiveness (+18.4%).
- automaticRed Team Test (ART):describeautomaticThe process and results of the Red Team Test (ART).Gemini 2.5 Flash and Pro maintainpowerfulWhile ensuring safety, it has become the most helpful model to date.
- Memory and Privacy:analyzeGemini 2.5 Detectable memory rate and privacy risks of the model, discoveredGemini The memory rate of model 2.5 was significantly lower than that of previous models, and no output containing personal privacy information was found.
- Security assessment and cutting-edge security framework: Description ofGemini 2.5 Pro conducts a comprehensive security assessment, including CBRN (chemical, biological, radiological, and nuclear information risks), cybersecurity,Machine LearningAssessments in areas such as research and development and deceptive alignment.
- External security testingDescribe the results of the external security testing plan, including the results of the external security testing program.Gemini The 2.5 Pro (Preview 05-06) assessment focuses on autonomous system risks, cybersecurity risks, CBRN risks, and social risks.
Get Google'sGemini 2.5 Model Technology White PaperOriginal PDF file, scan the QR code to follow and reply: 20250621
CodeFly, the world's first application to support the generation of Huawei HarmonyOS applications.AI Agent
In-depth experience of JinlingAICSDN officially launched its financial investment research platform.AI Agent