LPM 1.0 - An AI video generation model launched by Cai Haoyu of miHoYo.
LPM 1.0 (Large Performance Model) is a 17B-parameter video character performance generation model launched by Anuttacon (Cai Haoyu AI Company), which supports real-time full-duplex audio and video dialogue.
What is LPM 1.0?
LPM 1.0 (Large Performance Model) is a 17-parameter video character performance generation model launched by Anuttacon (Cai Haoyu AI Company), supporting real-time full-duplex audio and video dialogue. The model can transform a single image into a digital human that can speak, listen, react, and display subtle micro-expressions, maintaining consistent identity for an unlimited duration. LPM 1.0 is suitable as a general-purpose visual engine for scenarios such as AI dialogue, virtual live streaming, and game NPCs.
Main functions of LPM 1.0
- Real-time full-duplex dialogueIt supports real-time interaction where both parties can speak and listen simultaneously. Both parties can interrupt at any time, and the model can generate natural reactions such as pauses before responding and eye shifts.
- Unlimited duration, consistent identityBased on image input, it maintains the stability of details such as character appearance, teeth, facial texture, and profile contours in long videos of several hours, without the phenomenon of "becoming more and more distorted as it is generated".
- Three-mode controlThe character's performance is controlled through a combination of text (to control actions/expressions), audio (to drive lip movements/rhythm), and reference images (to maintain identity).
- Zero-shot generalizationIt supports any style, including realistic humans, 2D anime, 3D game characters, and non-human creatures, without the need for fine-tuning for specific fields.
- Emotional performanceThe model can generate subtle micro-expressions such as hesitation, thinking, and breathing rhythm, and supports lip-syncing with melody when singing.
Technical principles of LPM 1.0
- Data buildingBy strictly filtering the quality (retention rate <10%), we remove editing traces, beauty filters and other defects. We use an improved LR-ASD model to label the speaking/listening/idle state of each frame and achieve audio separation. At the same time, we construct multi-granularity identity reference conditions for global appearance, multi-view body and facial expressions to form a large-scale multimodal dataset.
- Base LPMBased on a 14B image-to-video pre-trained model, a 17B diffusion Transformer is formed by adding 3B parameters to interleaved audio attention blocks. This jointly learns speech-driven dynamics, listening responses, text control, and multi-reference identity preservation, and trains over 17 trillion multimodal tokens to achieve high-quality character performance generation.
- Online LPMThe four-stage autoregressive distillation course transforms Base LPM into a causal streaming generator. The Backbone-Refiner architecture preserves the time-series latent variable trajectories and restores high-fidelity details, enabling low-latency real-time inference and infinite-length identity consistency generation.
- System ArchitectureIt is plug-and-play compatible with the A2A audio model, and it cycles through the three states of listening, speaking, and idle to generate corresponding video streams in real time.
How to use LPM 1.0
LPM 1.0 is currently only for academic exchange and is not open to the public.
LPM 1.0 project address
- Project official websitehttps://large-performance-model.github.io/
- arXiv technical paper: https://arxiv.org/pdf/2604.07823
Key information and usage requirements of LPM 1.0
- definitionAnuttacon (Cai Haoyu AI Company) launched the 17B parameter video role performance model (Large Performance Model), which focuses on single-person full-duplex audio and video dialogue scenarios and can transform a single image into a digital human that can speak, listen, and react in real time.
- Core CompetenciesReal-time full-duplex dialogue (supports interruption), unlimited duration of identity consistency (long-term stable appearance/expression), trimodal control (text + audio + image), zero-sample generalization (supports realistic/animation/3D/non-human creatures), and delicate emotional performance (micro-expressions/breathing rhythm).
- technical routeThe Base LPM (17B Diffusion Transformer) is trained on a strictly filtered multimodal dataset, and then distilled into an Online LPM (Causal Streaming Architecture) through a four-stage process. Real-time generation is achieved using a Backbone-Refiner design.
- Application scenariosA general-purpose visual engine for dialogue agents, virtual live streaming, game NPCs, AI education tutors, and game companions.
- Current status:Not open to the publicThere are no model weights, source code, online demos, APIs, or any products. The project page is for academic exchange only.
The core advantages of LPM 1.0
- Solving the three dilemmas of performanceThe industry's first video generation model that simultaneously achieves high expressiveness, real-time reasoning, and long-term identity stability, breaking through the limitation of traditional models that can only take into account two of these aspects.
- Full-duplex real-time interactionIt supports true real-time dialogue, seamlessly switching between speaking and listening modes. Both parties can speak simultaneously and interrupt at any time. It has low response latency and features natural pauses, eye shifts, and other micro-reactions.
- Unlimited duration, consistent identityThe streaming architecture keeps the details of the character's appearance, teeth, and facial features stable over hours of video, preventing the identity drift that occurs with other models (such as Kling-Avatar 2.0/OmniHuman 1.5 limited to 30 seconds) over time.
- Natural listening behaviorThe model can generate realistic listening responses (nodding, eyebrow movement, eye contact), filling the gap in existing models that only focus on "speaking" and ignore "listening".
- Zero-shot generalizationThe model requires no fine-tuning and can support any style, including realistic humans, 2D anime, 3D game characters, and non-human creatures, possessing extremely strong character adaptability.
- SOTA performanceIt leads across the board on the first interactive character performance benchmark, LPM-Bench. In human evaluation, the preference rates for Kling-Avatar-2 and OmniHuman-1.5 in the 720P version were 64.3% and 42.5%, respectively.
Comparison of LPM 1.0 with similar competing products
| Comparison Dimensions | LPM 1.0 | Kling-Avatar 2.0 | OmniHuman-1.5 |
|---|---|---|---|
| Duration limit | Unlimited durationLong-term stable identity | Maximum 30 seconds | Maximum 30 seconds |
| Interaction mode | Full-duplex real-time(You can speak/listen/interrupt simultaneously) | One-way speech generation | One-way speech generation |
| Listening skills | Native support(Real-time reaction, nodding, eye contact) | Not supported | Not supported |
| Identity stability | Maintain consistency for several hours | It may drift over time. | It may drift over time. |
| Manual assessment | benchmark | 64.3%Users prefer LPM | 42.5%Users prefer LPM |
Application scenarios of LPM 1.0
- Conversational AI AgentIt gives AI assistants a concrete human visual presence, supports face-to-face real interaction, and can be used for customer support, virtual assistants, and digital humans.
- Interactive NPCs and Game CharactersCreate open-world NPCs with contextual dialogue, listening behavior, and emotionally coherent body language, enabling interactive storytelling without the need for separate motion capture.
- Live streaming and virtual hostingReal-time virtual streaming media can maintain identity consistency and visual quality even with hours of live streaming and sub-second latency, and supports 24/7 broadcasting.
- Education and Personalized TutoringAI tutors possess a continuous visual presence, maintaining a consistent identity throughout extended teaching sessions and seamlessly transitioning from enthusiastic explanation to focused listening.
- Game CompanionReal-time AI companions respond to gameplay with contextual comments, emotional encouragement, and natural facial expressions, adding a social interaction experience to single-player games.