Qwen3-235B-A22B-Thinking-2507 - Alibaba's latest inference model
Qwen3-235B-A22B-Thinking-2507 is Alibaba's most powerful open-source inference model globally. Based on a sparse hybrid expert (MoE) architecture with 235 billion parameters, it activates 22 billion parameters per activation and has 94 Transformer layers...
What is Qwen3-235B-A22B-Thinking-2507?
Qwen3-235B-A22B-Thinking-2507 is Alibaba's most powerful open-source inference model globally. Based on a sparse hybrid expert (MoE) architecture with 235 billion parameters, activating 22 billion parameters per cycle, it features a 94-layer Transformer network and 128 expert nodes. Designed for complex inference tasks, the model supports 256K native context processing capabilities and can handle long texts and deep inference chains. In terms of performance, Qwen3-235B-A22B-Thinking-2507 significantly improves core capabilities such as logical reasoning, mathematics, scientific analysis, and programming. In particular, it has achieved the best results among open-source models globally in benchmark tests such as AIME25 (mathematics) and LiveCodeBench v6 (programming), surpassing some closed-source models. It also performs well in general tasks such as knowledge acquisition, creative writing, and multilingual capabilities.
The model is licensed under the Apache 2.0 open-source license, free for commercial use. Users can experience and download it through QwenChat, Moda Community, or Hugging Face. The pricing is $0.70 per million input tokens and $8.40 per million output tokens.
Main functions of Qwen3-235B-A22B-Thinking-2507
-
Logical reasoningIt excels in logical reasoning tasks and is able to handle complex multi-step reasoning problems.
-
Mathematical operationsSignificant improvement in mathematical ability, especially in high-difficulty math tests such as AIME25, where it achieved the best results for open-source models.
-
Scientific AnalysisIt can handle complex scientific problems and provide accurate analysis and solutions.
-
Code generationIt can generate high-quality code and supports multiple programming languages.
-
Code optimizationIt helps developers optimize existing code and improve code efficiency.
-
Debugging supportProvides code debugging suggestions to help developers quickly locate and resolve problems.
-
256K context supportIt natively supports long text processing up to 256K, can handle extremely long contexts, and is suitable for complex document analysis and long conversations.
-
Deep inference chainIt automatically enables multi-step inference without requiring users to manually switch modes, making it suitable for tasks that require in-depth analysis.
-
Multilingual dialogueIt supports dialogue and text generation in multiple languages, meeting the needs of cross-language communication.
-
Instructions followedIt can accurately understand and execute user instructions, and generate high-quality text output.
-
Tool callIt supports integration with external tools to extend the functionality of the model.
Technical Principles of Qwen3-235B-A22B-Thinking-2507
- Sparse Hybrid Expert (MoE) ArchitectureQwen3-235B-A22B-Thinking-2507 adopts a sparse Mixture of Experts (MoE) architecture with a total of 235 billion parameters, activating 22 billion parameters per inference. This architecture contains 128 expert nodes, with each token dynamically activating 8 experts, balancing computational efficiency and model capability.
- Autoregressive Transformer structureThe model is based on an autoregressive Transformer architecture with 94 Transformer layers, supporting modeling of extremely long sequences and natively supporting 256K context lengths. This enables the model to handle complex long text tasks.
- Inference mode optimizationQwen3-235B-A22B-Thinking-2507 is designed specifically for deep reasoning scenarios and forces users into reasoning mode by default. It performs exceptionally well in fields requiring specialized knowledge, such as logical reasoning, mathematical operations, scientific analysis, programming, and academic assessment.
- Training and optimizationThe model further improves performance through a two-stage pre-training and post-training paradigm. In multiple benchmark tests, such as AIME25 (mathematics) and LiveCodeBench (programming), the model has achieved the best results among open-source models worldwide.
- Dynamic activation mechanismThe dynamic activation mechanism in the MoE architecture allows the model to dynamically select expert nodes based on task complexity during inference.
Project address for Qwen3-235B-A22B-Thinking-2507
- HuggingFace model libraryhttps://huggingface.co/Qwen/Qwen3-235B-A22B-Thinking-2507
Application scenarios of Qwen3-235B-A22B-Thinking-2507
-
Code generation and optimizationIt can generate high-quality code and help developers optimize existing code.
-
Creative WritingThey excel in creative writing, story creation, and copywriting, and can provide rich creative ideas and detailed concepts.
-
Academic writingIt can assist in writing academic papers and literature reviews, providing professional analysis and suggestions.
-
Research Design: To help design research plans and provide scientific and reasonable suggestions.