Textoon is an innovative project launched by Alibaba Group's Tongyi Lab, the first method to generate Live2D format 2D cartoon characters based on text prompts. Based on advanced language and visual models, it can generate diverse...
Textoon is an innovative project launched by Alibaba Group's Tongyi Lab, the first method to generate Live2D format 2D cartoon characters based on text prompts. Based on advanced language and visual models, it can generate diverse...
Ziyue-o1 is the first inference model in China released by NetEase Youdao that outputs step-by-step explanations. The model uses a 14B lightweight architecture, designed specifically for consumer-grade graphics cards, and can run stably on devices with low video memory. Through thought chain technology, it simulates...
Doubao Big Model 1.5 is the latest version of the big model released by ByteDance. It adopts a large-scale sparse MoE architecture, equivalent to the performance of a Dense model with 7 times the activation parameters, and achieves a high overall score across multiple evaluation criteria including knowledge, code, reasoning, and Chinese language performance...
OmniManip is a general-purpose robot manipulation framework developed by the joint laboratory of Peking University and Zhiyuan Robotics. By combining the high-level reasoning capabilities of a visual language model (VLM) with precise 3D manipulation capabilities, it enables robots to perform non-linear maneuvers...
WebWalker is a tool developed by Alibaba's Natural Language Processing team for evaluating and improving the performance of large language models (LLMs) in web browsing tasks. By simulating web navigation tasks, it helps models better handle long webpages...
VideoChat-Flash is a multimodal large language model (MLLM) for long video modeling, jointly developed by the Shanghai Artificial Intelligence Laboratory and Nanjing University, among other institutions. The model efficiently processes long videos using hierarchical compression technology (HiCo)...
EmoLLM is a large-scale language model focused on mental health support, providing users with emotional guidance and psychological support through multimodal emotion understanding. It combines various data formats such as text, images, and video, based on advanced multi-view perspectives...
Step-Video V2 is an upgraded video generation model released by Shanghai Jieyue Xingchen Intelligent Technology. This version features optimizations and innovations in several core technology areas, employing a VAE model with a higher compression ratio and deeply optimized DiT...
UI-TARS is a next-generation native graphical user interface (GUI) proxy model launched by ByteDance, enabling automated interaction with desktop, mobile devices, and web interfaces through natural language. It possesses powerful perception, reasoning, action, and...
EMO2 (End-Effector Guided Audio-Driven Avatar Video Generation) is an audio-driven avatar video generation technology developed by Alibaba Research Institute of Intelligent Computing. Its full name is "End-Effector Guided Audio-Driven Avatar Video Generation..."
PaSa is an AI agent for academic paper retrieval developed by ByteDance Research, based on reinforcement learning. It can mimic the behavior of human researchers, automatically calling search engines, browsing relevant papers, and tracking citations...
Baichuan-M1-preview is China's first full-scenario deep thinking model launched by Baichuan Intelligence. The model possesses reasoning capabilities in three major domains: language, vision, and search, and has performed excellently in multiple authoritative evaluations in mathematics, code, and other fields...
TokenVerse is a multi-concept personalized image generation method based on a pre-trained text-to-image diffusion model. It can decouple complex visual elements and attributes from a single image and seamlessly combine concepts extracted from multiple images to generate new images. ...
Baichuan-M1-14B is the industry's first open-source medical augmented model launched by Baichuan Intelligent. Its medical capabilities surpass those of the larger parameter-rich Qwen2.5-72B and are comparable to the o1-mini. It is specifically optimized for medical scenarios and possesses powerful...
CogVideoX-2 is an open-source text-to-video generation model from Zhipu AI. Based on an advanced 3D variational autoencoder (VAE), it compresses video data to 2% of its original size, reducing resource usage while ensuring the coherence between video frames...
CogView-4 is a text-to-image generation model developed by Zhipu AI. Based on the Transformer architecture, it's a diffusion model used to generate high-quality images. By optimizing parameter scaling and fine-tuning the dataset using high-quality images, it can generate...
LLMware is a unified framework designed for enterprise applications, suitable for building RAG (Retrieval-Augmented Generation) processes based on small, specialized models. LLMware supports private deployments and can be securely integrated with enterprise systems...
FilmAgent is a virtual film production tool developed by a research team at Harbin Institute of Technology (Shenzhen) based on a multi-agent collaborative framework. It automates the end-to-end film production process in a virtual 3D space, simulating traditional film production...
Whisper Input is an open-source voice input tool developed using Python and OpenAI's Whisper model. It allows for simple keyboard shortcuts (such as pressing the Option key to start recording and releasing it to stop recording) to...
Fast3R is a novel multi-view 3D reconstruction method proposed by researchers at Meta and the University of Michigan. Based on the Transformer architecture, it can process more than 1000 images in a single forward propagation process, achieving efficient and scalable 3D reconstruction...
Tarsier2 is an advanced large-scale visual language model (LVLM) launched by ByteDance. It generates detailed and accurate video descriptions and performs exceptionally well in various video understanding tasks. The model achieves performance improvements through three key upgrades...
VideoLLaMA3 is a cutting-edge multimodal foundational model open-sourced by Alibaba, focusing on image and video understanding. Based on the Qwen 2.5 architecture, it combines advanced visual encoders (such as SigLip) and powerful language generation capabilities...
Baichuan-Omni-1.5 is an open-source, full-modal model from Baichuan Intelligence. It supports full-modal understanding of text, images, audio, and video, and possesses bimodal generation capabilities for both text and audio. The model excels in visual, speech, and multimodal streaming...
TeleAI-t1-preview is a "Complex Reasoning Model" released by the Artificial Intelligence Research Institute of China Telecom, possessing powerful logical reasoning and mathematical derivation capabilities. Through reinforcement learning training methods and the introduction of thinking paradigms such as exploration and reflection,...