Open-Sora 2.0 is a brand-new open-source, state-of-the-art (SOTA) video generation model launched by Luchen Technology. Open-Sora 2.0 successfully trained a commercial-grade model with 11 parameters using $200,000 (224 GPUs)...
Open-Sora 2.0 is a brand-new open-source, state-of-the-art (SOTA) video generation model launched by Luchen Technology. Open-Sora 2.0 successfully trained a commercial-grade model with 11 parameters using $200,000 (224 GPUs)...
Gemini Robotics is a robotics project launched by Google DeepMind, based on Gemini 2.0, which brings the capabilities of large-scale multimodal models to the physical world. The project includes two main models: Gemini Robotics-ER and...
PP-TableMagic is a high-performance table recognition tool developed by the Baidu PaddlePaddle team. It's used to extract structured table information from images, convert it into HTML and other formats, and then perform further data processing and analysis. PP-TableMagic...
Gemini 2.0 Flash is a multimodal AI model from Google that combines text understanding and image generation capabilities. It generates high-quality images from natural language input, supports multi-turn conversational image editing, and maintains contextual coherence...
TokenSwift is an ultra-long text generation acceleration framework developed by the Beijing General Artificial Intelligence Research Institute team. It can generate text containing 100,000 tokens within 90 minutes, a 3x speed improvement compared to the nearly 5 hours required by traditional autoregressive models.
MIDI (Multi-Instance Diffusion for Single Image to 3D Scene Generation) is an advanced 3D scene generation technology that can convert a single image into a high-fidelity 3D scene in a short time. Through...
Evolving Agents is a production-grade framework for creating, managing, and evolving AI agents. Evolving Agents supports communication and collaboration between intelligent agents, evolving based on semantic understanding needs and past experience to effectively solve...
MT-MegatronLM is an open-source hybrid parallel training framework developed by Moore's Threads for full-featured GPUs, primarily used for efficiently training large-scale language models. It supports dense models, multimodal models, and MoE (Hybrid Experts)...
APB (Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUs) is a distributed long-context inference protocol proposed by Tsinghua University and other institutions...
Botgroup.chat is a multi-user AI chat application built on React and Cloudflare Pages. It supports multiple AI characters participating in a conversation simultaneously, providing a group chat-like interactive experience. Users can customize the personalities of the AI characters...
MT-TransformerEngine is a high-efficiency training and inference optimization framework open-sourced by Moore's Threads, specifically designed for Transformer models. The framework fully leverages Moore's Threads' full-featured GP... through operator fusion, parallel acceleration, and other techniques.
Chitu is a high-performance large-model inference engine jointly open-sourced by the Institute of High-Performance Computing at Tsinghua University and Tsinghua Jizhi. It is specifically designed to solve the high cost and low efficiency problems of large models during the inference stage and has powerful hardware adaptability...
Open-LLM-VTuber is an open-source, cross-platform voice-interactive AI companion project. It supports real-time voice dialogue and visual perception, features a vivid Live2D animated avatar, can run completely offline, and protects privacy. Users can use it as...
MetaStone-L1-7B is a lightweight inference model in the MetaStone series, designed to improve performance for complex downstream tasks. It achieves state-of-the-art performance among parallel models in core inference benchmarks such as math and code (SO...
Wenxin Big Model 4.5 is Baidu's latest generation and first native multimodal big model, officially released. It boasts significant improvements in multimodal understanding, text processing, and logical reasoning, outperforming GPT4.5 in multiple tests. The model is now available on Baidu AI...
Wenxin Big Model X1 is a deep thinking model launched by Baidu. It possesses "long thought chains" and excels in Chinese knowledge-based question answering, literary creation, and logical reasoning. X1 adds multimodal capabilities, enabling it to understand and generate images, and call upon tools to generate...
MM-Eureka is a multimodal reasoning model jointly developed by researchers from the Shanghai Artificial Intelligence Laboratory, Shanghai Institute of Innovation, Shanghai Jiao Tong University, and the University of Hong Kong. The model utilizes rule-based large-scale reinforcement learning (RL) to...
Command A is Cohere's latest generative AI model, designed specifically for enterprise applications. Command A's core advantages lie in its high performance and low hardware cost, enabling efficient deployment on two GPUs, compared to other similar models...
AudioX is a unified diffusion transformer model jointly proposed by the Hong Kong University of Science and Technology and Dark Side of the Moon, specifically designed for generating audio and music from arbitrary content. The model can handle various input modalities, including text, video, images, music, and...
MedRAG is a medical diagnostic model proposed by a research team at Nanyang Technological University. It enhances the diagnostic capabilities of Large Language Models (LLMs) by incorporating knowledge graph reasoning. The model constructs a four-layer fine-grained diagnostic knowledge graph, enabling precise classification of...
I2V3D is an innovative image-to-video generation framework developed by City University of Hong Kong and Microsoft GenAI. It supports the conversion of still images into dynamic videos, achieving precise animation control based on 3D geometry guidance. I2V3D combines traditional computer graphics...
OpenBioMed is an open-source platform jointly launched by the Institute for Intelligent Industry at Tsinghua University (AIR) and SMTH Blog, focusing on AI-driven biomedical research. It is a multimodal representation learning toolkit capable of processing molecules, proteins, ...
AMISS is an open-source low-code front-end framework from Baidu. It quickly generates various backend pages based on simple JSON configuration, eliminating the need for complex front-end code. AMISS supports forms, tables, charts, CRUD operations, and provides rich functionality...
Mistral Small 3.1 is an open-source multimodal AI model from Mistral AI, featuring 24 billion parameters and released under the Apache 2.0 license. It excels in text and multimodal tasks, supporting datasets up to 128k bytes...