FastDeploy is a high-performance inference and deployment tool developed by Baidu based on the PaddlePaddle framework, specifically designed for Large Language Models (LLMs) and Visual Language Models (VLMs). FastDeploy supports various hardware...
FastDeploy is a high-performance inference and deployment tool developed by Baidu based on the PaddlePaddle framework, specifically designed for Large Language Models (LLMs) and Visual Language Models (VLMs). FastDeploy supports various hardware...
DragonV2.1 (DragonV2.1Neural) is Microsoft's latest zero-shot text-to-speech (TTS) model. Based on the Transformer architecture, the model supports multiple languages and zero-shot speech cloning, requiring only 5-90 seconds of speech...
Wuhr AI Ops is an intelligent operations and maintenance management platform that simplifies complex operations and maintenance tasks through AI technology. The platform integrates a multimodal AI assistant, supports natural language interaction for executing operations and maintenance commands, and can switch between Kubernetes cluster and Linux system commands with a single click...
Skywork MindLink is an open-source inference model launched by Kunlun Wanwei. It features an adaptive inference mechanism, flexibly switching inference modes according to task complexity. It quickly generates inference models for simple tasks and performs deep inference for complex tasks, balancing efficiency...
ScreenCoder is an open-source intelligent UI screenshot-to-code system that supports quickly converting any design screenshot into clean, editable HTML/CSS code. ScreenCoder uses a modular multi-agent architecture combined with visual understanding...
RedOne is Xiaohongshu's first customized Large Language Model (LLM) for the social networking service (SNS) domain. The model employs a three-stage training strategy, incorporating social and cultural knowledge to enhance multi-task capabilities and align with platform standards...
Windows-MCP is a lightweight, open-source tool for integrating AI agents with Windows systems. As an MCP server, Windows-MCP allows Large Language Models (LLMs) to directly interact with Windows, enabling file browsing, application access, and more.
MCP Server Chart is a visualization chart generation tool developed by the AntV team. Based on the Model Context Protocol (MCP), the tool supports over 25 types of visualization charts, including common statistical charts (such as...).
OAgents is an open-source foundational agent framework launched by OPPO PersonalAI Lab. Based on standardized evaluation protocols and modular design, the framework promotes research in agent frameworks. OAgents analyzes key AI features based on empirical system research...
WorldVLA is an autoregressive action-world model jointly developed by Alibaba DAMO Academy and Zhejiang University. The model integrates the Visual-Language-Action (VLA) model with a world model into a single framework. The model is based on action and image processing...
AnimaX is a high-efficiency 3D animation generation framework developed by Beijing University of Aeronautics and Astronautics in collaboration with Tsinghua University, the University of Hong Kong, and others. It combines the action priors of a video diffusion model with a skeleton-based animation structure. The framework can generate animations from videos...
Ovis-U1 is a multimodal unified model developed by Alibaba Group's Ovis team, boasting 3 billion parameters. The model integrates three core capabilities: multimodal understanding, text-to-image generation, and image editing, based on an advanced architecture and collaborative unification...
Deep Video Discovery (DVD) is a deep video discovery agent from Microsoft, designed specifically for understanding and analyzing long videos. Deep Video Discovery segments long videos into multiple shorter segments, based on large-scale language processing...
FairyGen is an animated story video generation framework developed by universities in the Greater Bay Area. It supports generating animated story videos with coherent narratives and a consistent style, starting from a single hand-drawn character sketch. The framework leverages a multimodal large-scale language model (...
OmniGen2 is an open-source multimodal generative model developed by the Beijing Academy of Artificial Intelligence. It can generate high-quality images based on text prompts and supports instruction-guided image editing, such as modifying backgrounds or facial features. OmniGen2...
Qwen-TTS is a speech synthesis model developed by Alibaba Tongyi, characterized by its natural, stable, and fast performance. The model can output high-quality audio based on text and timbre parameters, supporting the synthesis of Chinese, English, and dialects such as Beijing Mandarin, Shanghainese, etc.
Speaker is a free and open-source AI meeting assistant that automates meeting recording transcription, content summarization, and intelligent question-and-answer while ensuring absolute data privacy. Speaker can run without an internet connection, and all data processing is handled remotely...
Goedel-Prover-V2 is an open-source theorem prover jointly developed by top institutions such as Princeton University, Tsinghua University, and NVIDIA. Goedel-Prover-V2 utilizes hierarchical data synthesis, verifier-guided self-correction, and model...
MirageLSD, developed by the Decart AI team, is the world's first Live-Stream Diffusion AI video model. It enables real-time video generation of unlimited duration with latency as low as 40 milliseconds and supports...
ChatFlow is an open-source, easy-to-use workflow engine that combines user-designed, high-quality workflows with AI-generated capabilities. ChatFlow supports visual components and automated execution, helping developers quickly generate code,...
Fogsight is an animation-generating agent driven by a large language model (LLM). Users input abstract concepts or words, and it generates high-quality, vivid animations. Its core functionality includes "concept-as-image," which transforms input topics into...
OpenBB is an open-source financial platform that provides powerful investment research tools for individuals and businesses. The platform integrates various financial data, including stocks, options, cryptocurrencies, forex, macroeconomics, and fixed income, and supports Python...
OpenReasoning-Nemotron is a series of powerful large language models (LLMs) open sourced by NVIDIA. It is distilled from the DeepSeek R1 0528 model and has parameter sizes of 1.5B, 7B, 14B and 32B.
Seed-X is an open-source multilingual translation model developed by ByteDance's Seed team. It boasts 7 billion parameters and supports bidirectional translation in 28 languages. Seed-X combines high-quality multilingual data pre-training, instruction fine-tuning, and reinforcement learning...