T2V-Turbo - Google's open-source text-to-video generation model
T2V-Turbo is an advanced text-to-video generation model developed by researchers from Google, UC Santa Barbara, and the University of Waterloo...
T2V-Turbo is an advanced text-to-video generation model developed by researchers from Google, UC Santa Barbara, and the University of Waterloo...
FLUX.1-Turbo-Alpha is an 8-step distillation LoRa model trained by the Alimama creative team based on the FLUX.1-dev model. Based on multi-head discriminator technology, it improves the quality of generated images and supports text-to-image generation and inpainting control networks...
OpenR is a full-chain training framework jointly open-sourced by University College London (UCL), Shanghai Jiao Tong University, University of Liverpool, Hong Kong University of Science and Technology (Guangzhou), and Westlake University. It aims to improve the performance of large language models (LLMs)...
Agent-S is an innovative agent framework designed to automate human-computer interaction based on a graphical user interface (GUI). Agent-S simulates human operation, allowing direct interaction with the computer using a mouse and keyboard to handle complex tasks...
Adobe Firefly is a suite of creative generative AI models from Adobe designed to help users expand their innate creativity. These models are integrated into Adobe's flagship applications and Adobe Stock, supporting...
Augmented Physics is an innovative educational tool that uses integrated machine learning technology to transform static diagrams in physics textbooks into interactive and embedded physics simulations. The tool is based on advanced computer vision technology...
podlm-public is an open-source AI podcast tool designed to create a Chinese alternative to NotebookLM, specifically for converting arbitrary URLs into podcast content and then pushing it to the Xiaoyuzhou platform. The project is based on advanced AI technology...
Yi-Lightning, the latest flagship model released by Zero-One-Way Technology Co., Ltd., has achieved remarkable results on the internationally authoritative blind benchmarking list LMSYS, surpassing OpenAI's GPT-4o-2024-05-13 and Anthropic C...
FunASR is an open-source speech recognition toolkit from Alibaba DAMO Academy, providing features including speech recognition (ASR), voice activity detection (VAD), punctuation recovery, language modeling, speaker verification, speaker segmentation, and multi-speaker ASR...
CleanS2S is a prototype of a streaming speech-to-speech (S2S) interactive intelligent agent, providing a high-quality, real-time voice interaction experience. The CleanS2S project is implemented using a single file, simplifying the configuration and understanding process, making it convenient for users and researchers...
Hallo2 is an audio-driven video generation model jointly developed by Fudan University, Baidu, and Nanjing University. It combines a single reference image with several minutes of audio input, adjusting portrait expressions based on optional text prompts...
Model Judge is an online AI model evaluation platform built on Next.js. Users input questions and select multiple AI models to test, helping them quickly identify the most suitable AI model for their needs. The platform's unique feature is its...
AgentStack is an open-source tool designed to help developers quickly build AI agent projects. Based on providing a pre-configured template and integrating popular agent frameworks and large language model (LLM) providers, it simplifies the creation process from scratch...
Marco is Alibaba International's latest large-scale commercial translation model, supporting 15 major global languages, including Chinese, English, Japanese, Korean, Spanish, and French. It surpasses competitors such as Google Translate, DeepL, and GPT-4 in the BLEU benchmark...
Ministral 3B and 8B are two new AI mini-models launched by Mistral AI, designed specifically for on-device computing and edge applications. They offer strengths in knowledge, common sense, reasoning, function calls, and efficiency for categories with fewer than 1 billion parameters...
TANGO is an open-source framework jointly developed by the University of Tokyo and CyberAgent AI Lab, focusing on generating full-body gesture videos synchronized with the target's speech. Based on hierarchical audio motion embedding and diffusion interpolation networks, it synchronizes the target speech...
Nemotron-70B-Instruct is a large-scale language model released by NVIDIA. It utilizes a novel hybrid training method to improve the model's response quality and consistency when following instructions. The model combines Bradley-Ter...
SANA is a text-to-image generation framework jointly developed by NVIDIA, MIT, and Tsinghua University. It can efficiently generate high-resolution images up to 4096×4096. SANA is based on a deep compression autoencoder, linear expansion...
Chat2DB is an AI-driven database management and analysis tool based on natural language processing technology. It allows users to interact with the database using natural language, simplifying SQL code writing and database management. Chat2DB supports various...
IterComp is a text-to-image generation framework jointly developed by researchers from Tsinghua University, Peking University, LibAI Lab, University of Science and Technology of China, Oxford University, and Princeton University. It is based on an iterative feedback learning mechanism...
LayerSkip is a technique used to accelerate the inference process of large language models (LLMs). Based on applying layer dropout and early exit loss during the training phase, it allows the model to exit more accurately from earlier layers during inference, without requiring...
Spirit LM is a multimodal language model developed by the Meta AI team that seamlessly mixes text and speech data. Spirit LM is based on a pre-trained text language model, continuously trained on text and speech units...
Story-Adapter is a novel long-form story visualization framework that generates high-quality, subtly interactive sequences of story images while maintaining semantic consistency. It achieves this through an iterative approach, based on global reference cross-attention...
LOKI is a synthetic data detection benchmark jointly proposed by Sun Yat-sen University and Shanghai AI Lab. It aims to comprehensively evaluate the ability of large-scale multimodal models (LMMs) to recognize synthetic data across multiple modalities, including video, images, 3D, text, and audio.