AB
AiBoss
project

Tencent Hunyuan Turbo S - Tencent's next-generation fast thinking model

Tencent's Hunyuan Turbo S is a new generation of fast-thinking model launched by Tencent. The model adopts an innovative Hybrid-Mamba-Transformer fusion architecture, effectively reducing the computational complexity of traditional Transformers and minimizing KV-Cache...

What is Tencent Hunyuan Turbo S?

Tencent's Turbo S is a next-generation fast-thinking model launched by Tencent. The model employs an innovative Hybrid-Mamba-Transformer fusion architecture, effectively reducing the computational complexity of traditional Transformers, minimizing KV-Cache usage, and significantly improving training and inference efficiency. As the industry's first lossless application of the Mamba architecture to ultra-large-scale MoE models, Turbo S performs exceptionally well in multiple domains, including knowledge, mathematics, and reasoning, comparable to leading models such as DeepSeek V3 and GPT-4o.

The core advantage of Hunyuan Turbo S lies in its rapid response, achieving "instant reply," doubling the speed of speech and reducing the latency of the first character by 44%. It performs exceptionally well in short thought chain tasks (such as mathematics, coding, and logical reasoning), while also incorporating the long thought chain capabilities of the Hunyuan T1 slow thinking model, balancing stability and accuracy.

Main functions of Tencent Hunyuan Turbo S

  • rapid response capabilityThe Hunyuan Turbo S can achieve "instant response", doubling the speed of speech and reducing the latency of the first character by 44%, significantly improving the smoothness of interaction and user experience.
  • Multi-domain knowledge and reasoning abilityIt performs exceptionally well in multiple fields such as knowledge, mathematics, and logical reasoning, and is comparable to industry-leading models such as DeepSeek V3 and GPT-4o.
  • Content creation and multimodal supportIt supports high-quality literary creation, text summarization, multi-turn dialogue and other functions, and also has the multimodal ability to generate images from text.
  • Low deployment cost and high cost-effectivenessIt adopts a Hybrid-Mamba-Transformer integrated architecture, which reduces the computational complexity and deployment cost of traditional Transformers.

The technical principles of Tencent Hunyuan Turbo S

  • Advantages of the Mamba architectureThe Mamba architecture is based on the State Space Model (SSM) and, by introducing a selective mechanism, can efficiently process long sequence data. It performs exceptionally well when processing long text, while significantly reducing computational complexity and KV-Cache usage.
  • Retention of the Transformer architectureThe Transformer architecture excels at capturing complex contextual relationships, and Hunyuan Turbo S retains this advantage. At the same time, by integrating the Mamba architecture, it breaks through the bottleneck of traditional Transformer in long text processing and inference costs.
  • Optimization of the MoE modelMixed Turbo S is the first successful application of the Mamba architecture losslessly to ultra-large-scale MoE (Mixture of Experts) models in the industry. It improves the model's memory and computational efficiency while reducing training and inference costs.
  • Integration of long and short thinking chainsWhile maintaining the rapid response (quick thinking) experience for humanities questions, Hunyuan Turbo S significantly improves the reasoning ability for science questions through its self-developed long thought chain data, thereby improving the overall performance of the model.

Performance of Tencent Hunyuan Turbo S

  • Knowledge domain:
    • In the MMLU benchmark test, the mixed-source Turbo S scored 89.5, slightly lower than DeepSeek V3's 88.5, but higher than other models.
    • In the MMLU-pro test, the Hunyuan Turbo S scored 79.0, outperforming GPT4o-0806 and Claude-3.5.
    • In the GPQA-diamond test, the Hunyuan Turbo S scored 57.5, outperforming the Llama3.1-405B and DeepSeek V3.
    • In the SimpleQA test, the mixed-source Turbo S scored 22.8, which is not as good as other models.
    • In the Chinese-SimpleQA test, the Hunyuan Turbo S scored 70.8, outperforming GPT4o-0806 and Claude-3.5.
  • Reasoning Domain:
    • In the BBH test, the Mixed Turbo S scored 92.2, outperforming all other models.
    • In the DROP test, the Hunyuan Turbo S scored 91.5, outperforming the GPT4o-0806 and Claude-3.5.
    • In the ZebraLogic test, the Mixed Turbo S scored 46.0, which is not as good as other models.
  • Mathematics:
    • In the MATH test, the Hunyuan Turbo S scored 89.7, outperforming the GPT4o-0806 and Claude-3.5.
    • In the AIME2024 test, the Hunyuan Turbo S scored 43.3, outperforming GPT4o-0806 and Claude-3.5.
  • Code Domain:
    • In the HumanEval test, the Hunyuan Turbo S scored 91.0, outperforming GPT4o-0806 and Claude-3.5.
    • In the LiveCodeBench test, the Hunyuan Turbo S scored 32.0, which is not as good as other models.
  • Chinese Language Field:
    • In the C-Eval test, the Hunyuan Turbo S scored 90.9, outperforming GPT4o-0806 and Claude-3.5.
    • In the CMMLU test, the Hunyuan Turbo S scored 90.8, outperforming GPT4o-0806 and Claude-3.5.
  • Alignment Area:
    • In the ArenaHard test, the Hunyuan Turbo S scored 88.6, outperforming the GPT4o-0806 and Claude-3.5.
    • In the IF-Eval test, the Hunyuan Turbo S scored 88.6, outperforming GPT4o-0806 and Claude-3.5.

How to use Tencent Hunyuan Turbo S

  • Tencent Cloud Official WebsiteThe Hunyuan Turbo S model has been officially launched on the Tencent Cloud website, and developers and enterprise users can call the model via API.
  • Tencent YuanbaoThe model will be gradually rolled out in the Tencent Yuanbao APP. Users can select the "Hunyuan" model and turn off the deep thinking function to experience it.
  • Free trialStarting today, developers and enterprise users can enjoy a one-week free trial of Hunyuan Turbo S via API on Tencent Cloud. VisitTencent Hunyuan Turbos Model API Free Trial ApplicationPlease fill in your address in the application.
  • Future plansThe Hunyuan Turbo S will become the core foundation of Tencent's Hunyuan series of derivative models, providing basic capabilities for derivative models such as reasoning, long articles, and code.

Tencent Hunyuan Turbo S model pricing

  • Model pricingThe API call pricing for Hunyuan Turbo S is 0.8 yuan per million tokens for input and 2 yuan per million tokens for output.

Application scenarios of Tencent Hunyuan Turbo S

  • Daily conversationSuitable for scenarios such as rapid Q&A and intelligent customer service.
  • Code generation and logical reasoningIt performs exceptionally well in short thought chain tasks such as mathematics, code generation, and logical reasoning.
  • Content creationSupports high-quality text generation and text-to-image generation.