AB
AiBoss
project

WeLM - WeChat's self-developed large language model

WeLM (WeChat Language Model) is a large language model developed by WeChat, which has been iterated to V4. The model is specifically designed for the WeChat ecosystem, with approximately 80B-130B parameters. It adopts a sparse MoE architecture, emphasizing low cost and high efficiency.

What is WeLM?

WeLM (WeChat Language Model) is a large-scale language model developed by WeChat, now at version V4. Specifically designed for the WeChat ecosystem, the model has approximately 80-130 bytes of parameters and employs a sparse MoE architecture, emphasizing low cost and high efficiency. WeLM supports 128K long context and can naturally acquire ecosystem data from group chats, Moments, and other sources. It uses Hidden Decoding technology to hide the inference process, enabling rapid response. WeLM is not publicly available; it serves as the intelligent foundation for WeChat's AI Agent, Xiaowei.

Main functions of WeLM

  • Exclusive intelligent services within the WeChat ecosystemAs the underlying model of WeChat's AI Agent "Xiaowei", WeLM is deeply embedded in the WeChat ecosystem, supporting scenarios such as summarizing group chat messages, viewing Moments updates, replying to friends' messages, and understanding content from official accounts/video accounts.
  • Low-cost, high-efficiency reasoningIt adopts a highly sparse MoE (hybrid expert) architecture, combined with technologies such as GQA, KV-Mirror, and Multi-Token Prediction, which significantly reduces inference costs and supports high-frequency calls from WeChat's 1 billion+ daily active users.
  • Fast response (Hidden Decoding)Hidden Decoding technology hides the inference process, ensuring output quality while achieving low-latency response and avoiding long user wait times.
  • Long context understandingIt supports 128K ultra-long context, naturally acquiring ecological data such as group chat records, Moments, favorites, and likes, to achieve accurate intent recognition and personalized replies.
  • MultitaskingIt covers a variety of tasks such as content summarization, dialogue generation, information retrieval, and personalized recommendations, becoming the core engine of AI capabilities within WeChat.

How to use WeLM

  • Access the official portalSearch for "Xiaowei" on WeChat to find the access point.
  • Update WeChatUpgrade WeChat to the latest version to ensure you can receive "Xiaowei" beta push notifications or full rollout.
  • Find the entranceFind the "Xiaowei" AI Agent entry in the WeChat chat list, top search bar, or Discover page.
  • Initiate a dialogueClick "Xiaowei" to directly enter your questions and task requirements via text or voice.
  • Authorized dataAuthorize access to ecosystem data such as group chat history, Moments, and favorites as prompted to receive personalized, context-aware responses.
  • Execute the taskUse "Xiaowei" for tasks such as group chat summarizing, content recommendation, script generation, and information retrieval.
  • Optimize interactionBased on feedback from "Xiaowei", your instructions will be refined through multiple rounds of dialogue to improve the quality of the results.

WeLM's core advantages

  • WeChat ecosystem native integrationWeLM is a "dedicated mini-model" designed specifically for the WeChat ecosystem. It naturally acquires contextual data from group chats, Moments, official accounts, video accounts, and favorites, without requiring users to manually upload or authorize external tools, and can more accurately understand user intent.
  • Extreme cost controlIt adopts a highly sparse MoE architecture (not a super-dense model), combined with technologies such as GQA, KV-Mirror, and Multi-Token Prediction, which greatly reduces inference costs and supports high-frequency calls from WeChat's 1 billion+ daily active users.
  • Fast response experienceBy using Hidden Decoding technology to hide the inference process, it achieves low latency while ensuring output quality, avoids user waiting, and is suitable for WeChat instant messaging scenarios.
  • Extremely long contextual memoryIt supports 128K long context and performs excellently in long text tasks such as group chat summaries, Moments recaps, and personalized recommendations.
  • Stable and controllable closed architectureIt is not open to the public and focuses on serving internal WeChat scenarios, giving it a natural advantage in stability, security, and ecosystem consistency.

Comparison of WeLM's similar products

Dimension WeLM ByteDance
position WeChat ecosystem dedicated closed model ByteDance's universal AI assistant (open to all platforms)
Parameter size 80B-130B ("small model") A general-purpose large model with hundreds of billions of users
Architecture Highly sparse MoE, achieving extreme cost reduction Dense + MoE blend, pursuing general performance
Data fusion Natively acquire WeChat ecosystem data (group chats, Moments, etc.) Users need to manually upload or authorize external data.
Response strategy Hidden Decoding, a speed-oriented puzzle game. Supports displaying the reasoning process, allowing for in-depth thinking.
Openness Not open to the public, only available within WeChat. Independent Apps, Websites, and APIs are fully open.
Core Objectives Low cost, low latency, high stability, and deep ecosystem General capabilities, multimodal, and large-scale user coverage

Application scenarios of WeLM

  • Group chat intelligent summaryWeLM can automatically sort through massive amounts of historical messages in WeChat group chats, accurately extract key information, core decisions, and to-do items, helping users quickly grasp the dynamics of group chats.
  • WeChat Moments Interaction AssistantWeLM can intelligently generate appropriate and personalized comments or reply suggestions based on friends' Moments posts and the relationship between users and their friends.
  • Ecological Content Q&AWeLM can provide in-depth analysis of WeChat official account articles, video accounts, and other ecosystem content, and instantly answer users' questions about article details, video themes, and other aspects.
  • Personal memory retrievalWeLM supports cross-database interactions, including chat history, favorites, and likes, helping users quickly and accurately locate and retrieve important past information.
  • Intelligent script generationWeLM can combine the current chat context, friend relationships, and personal style to recommend appropriate and natural replies and communication techniques to users.