HiFiVFS (High Fidelity Video Face Swapping) is a high-fidelity video face-swapping framework developed by Tencent and VIVO. HiFiVFS is based on the Stable Video Diffusion (SVD) framework, using multi-frame input and time...
HiFiVFS (High Fidelity Video Face Swapping) is a high-fidelity video face-swapping framework developed by Tencent and VIVO. HiFiVFS is based on the Stable Video Diffusion (SVD) framework, using multi-frame input and time...
MVGenMaster is a multi-view diffusion model jointly developed by Fudan University, Alibaba DAMO Academy, and Hupan Lab. It's based on enhanced 3D prior processing for diverse novel view synthesis (NVS) tasks. The model is based on metric depth and camera...
360Zhinao2-7B is an upgraded version of 360's self-developed AI large model, 360 Zhinao 7B, encompassing basic models and chat models with various context lengths. The 360Zhinao2-7B model is a significant update following 360Zhinao1-7B, based on...
GeneMAN is a 3D human creation framework jointly developed by the Shanghai AI Lab, Peking University, Nanyang Technological University, and Shanghai Jiao Tong University. It can create high-fidelity 3D human models from a single image. The framework does not rely on parametric human models...
MagicDriveDiT is a novel video generation method based on the DiT architecture, jointly developed by the Chinese University of Hong Kong, the Hong Kong University of Science and Technology, Huawei Cloud, and Huawei Noah's Ark Lab. It is specifically designed for autonomous driving applications, achieving high resolution and long...
EfficientTAM is a lightweight video object segmentation and tracking model from Meta AI that addresses the high computational complexity of deploying the SAM 2 model on mobile devices. It is based on a simple, non-hierarchical Vision Transformer...
Amazon Nova is a next-generation family of AI foundational models from Amazon Web Services (AWS), offering industry-leading performance and cost-effectiveness. The family includes Amazon Nova Micro, specifically designed for text processing, and Amazon Nova for multimodal applications...
HunyuanVideo is an open-source video generation model from Tencent, boasting 13 billion parameters, making it one of the most parameter-rich open-source video models currently available. HunyuanVideo features robust physical simulation, high text semantic fidelity, motion consistency, and...
Lobe Vidol is an open-source digital human creation platform that allows everyone to easily create and interact with their own virtual idols. Lobe Vidol offers a smooth dialogue experience, background settings, a library of motion poses, an elegant user interface, and more...
GPT Academic is a feature-rich open-source project designed specifically for academic research and writing. It integrates one-click paper translation, source code parsing, internet information retrieval, LaTeX proofreading, and more...
Vanna is an open-source Python RAG (Retrieval-Augmented Generation) framework that helps users generate accurate SQL queries for their databases based on large language models (LLMs). Vanna uses a simple two-step process...
PersonaCraft is a personalized full-body image synthesis technology developed by Seoul National University in South Korea. Combining diffusion models and 3D human modeling, it can generate realistic, personalized full-body images of multiple people from a single reference image. PersonaCraft...
StableAnimator is an end-to-end high-quality identity-preserving video diffusion framework jointly developed by Fudan University, Microsoft Research Asia, Huya, and Carnegie Mellon University. StableAnimator can generate videos based on a reference image and...
I2V-01-Live is an image-to-video model launched by Conch AI, capable of converting static 2D images into dynamic videos. Based on deep learning technology, the model enhances the smoothness and vividness of movements, making the actions of people or objects more natural and...
Genie 2 is a new generation of large-scale base world models from DeepMind, capable of generating interactive 3D game worlds up to one minute long from a single image. Genie 2 can simulate complex dynamics such as object interactions, character animations, and physics effects...
Luma Photon is a next-generation image generation model from Luma AI, delivering ultra-high image quality and low-cost efficiency through an innovative architecture. Luma Photon supports personalized and creative image generation and can understand natural language...
TeleAI Video Generation Model is a video generation model launched by the AI Research Institute of China Telecom. It's based on a two-stage generation framework: first, it creates storyboard sketches based on text descriptions, and then generates the video based on these sketches. TeleAI Video Generation Model...
TPDM (Time Prediction Diffusion Model) is an image generation model jointly developed by the MAPLE Lab at Westlake University, Southern University of Science and Technology, Peking University, and the Institute of Advanced Technology at Westlake University. It can automatically...
ConsisID is a text-to-video (IPT2V) generation model developed by Peking University and Pengcheng Laboratory, among other institutions. It maintains the consistency of identity among individuals in a video based on frequency decomposition technology. The model uses a no-tuning (tuning-free) approach...
Perplexideez is a native AI assistant that enables users to quickly search for information on the web and in self-hosted applications. The Perplexideez project is based on the Postgres database, supports Ollam or OpenAI-compatible endpoints, and uses SearchXNG...
GenCast is a revolutionary AI weather forecasting model from DeepMind, based on diffusion modeling technology, providing global weather forecasts up to 15 days in advance. GenCast outperforms the world's leading medium-range weather forecasting systems in 97.2% of forecasting tasks...
FullStack Bench is a new code evaluation benchmark jointly launched by ByteDance's Doubao Big Model team and the M-A-P community, focusing on evaluating full-stack programming and multi-language programming skills. FullStack Bench covers more than 11 real-world programming languages...
Motion Prompting is a video generation technology jointly developed by Google DeepMind, the University of Michigan, and Brown University. It controls and guides the generation of video content based on motion trajectories. Motion...
Fish Speech 1.5 is a text-to-speech (TTS) model from Fish Audio, based on deep learning technologies such as Transformer, VITS, VQVAE, and GPT. Fish Speech 1.5 supports English, Japanese, Korean, ...