JoyAI-Echo - JD.com's open-source long audio/video generation framework
JoyAI-Echo is an open-source long-form video generation framework launched by JD.com, designed specifically for minute-level multi-shot story generation. The framework utilizes a cross-modal memory library, memory-driven post-training, Director Agent conversational editing, and lightweight...
What is JoyAI-Echo?
JoyAI-Echo is an open-source long-form video generation framework launched by JD.com, designed specifically for generating multi-shot stories at the minute level. Through four major technological innovations—cross-modal memory library, memory-driven post-training, Director Agent conversational editing, and lightweight real-time super-resolution—the framework addresses core pain points in long-form video generation, such as character face-changing, abrupt changes in voice, and slow generation. It achieves, for the first time, the generation of highly consistent, interactive, and high-definition long videos up to 5 minutes in length, marking JD.com's entry into the world's leading tier of long-form video generation.
Main functions of JoyAI-Echo
-
Minute-level multi-camera story generationSupports generating coherent, multi-camera long video sequences, up to 5 minutes long, from a single cue word JSON.
-
Cross-modal audio and video co-generationA single pipeline outputs video and audio simultaneously, ensuring audio-visual synchronization.
-
Paired cross-modal memory bank: Continuously save and call up character appearance features and speaker voice timbre during multi-camera generation to maintain story-level consistency.
-
DMD distillation short-step reasoningBy using distribution matching distillation technology, the generation rate is increased by approximately 7.5 times.
-
Director Agent (Conversational Editing)Users can interact with the director's assistant using natural language, and the script, characters, scenes, and shots are automatically broken down, supporting partial revisions without having to rerun the entire video.
-
Lightweight real-time super-resolutionSupports single-step super-resolution from 736×1280 to 1152×1920 or 1472×2560, maintaining high-definition output under streaming delay constraints.
The technical principles of JoyAI-Echo
- Cross-modal audio and video memoryJoyAI-Echo's core breakthrough lies in its built-in paired cross-modal memory bank, which binds and stores visual and audio memories through a slot-paired mechanism. During multi-shot generation, the memory bank continuously saves and retrieves the character's facial features, overall appearance, speaker's timbre, and audio-visual correspondence, ensuring that each new shot is generated based on the identity features of previous shots. This maintains story-level consistency throughout a 5-minute video, completely resolving issues of character facial changes and abrupt voice shifts.
- Memory-driven post-training and DMD distillation accelerationThe team introduced a memory-driven post-training workflow that combines Supervised Fine-Tuning (SFT), cross-modal RLHF, and Distribution Matching Distillation (DMD) techniques. Among them, DMD compresses the original multi-step diffusion inference into a few steps of inference, achieving an inference speedup of approximately 7.5 times while maintaining the generation quality, making the streaming generation of minute-long videos a practical application instead of a theoretical one.
- Director Agent Interaction ArchitectureThe framework introduces an intelligent director agent that automatically expands the user's natural language intent into structured scripts, shots, character descriptions, and scene descriptions, supporting a closed-loop workflow of planning, generation, review, and partial revision. Users can specify modifications through dialogue, and the agent only regenerates the problematic shots without rerunning the entire video, transforming static generation into dynamic collaboration.
- Lightweight real-time audio and video super-resolutionTo meet the high-definition requirements of professional content production, JoyAI-Echo is equipped with a single-step audio and video super-resolution module, which can sharpen the basic output of 736×1280 to 1152×1920 or 1472×2560 in real time under streaming latency constraints, ensuring that high-resolution output does not break the real-time performance of streaming generation.
How to use JoyAI-Echo
-
Cloning repository:
git clone https://github.com/jd-opensource/JoyAI-Echo.git -
Creating an environmentInstall dependencies using Python 3.11 + PyTorch 2.8 + CUDA 12.8 via conda or uv, and ensure...
ffmpegAvailable. -
Download model weightsDownload approximately 46GB from Hugging Face
echo-longvideo-release.safetensorsand approximately 24GBgemma-3-12bText encoder, placed incheckpoints/Table of contents. -
Write story promptsCreate a JSON file describing each shot in the order of character and subject, action and dialogue, style, camera movement, background, sound effects and background music.
-
Running inference:implement
python inference.pyAfter the model is loaded once, all prompt files are processed and output to [the appropriate server/system].inference_result/outputs/Table of contents.
JoyAI-Echo's core advantages
-
Ultra-long consistencyIn a 5-minute video, the character's identity, visual appearance, and voice timbre remain highly consistent, completely solving the problem of the same person becoming a different person while acting.
-
Rapid generationMemory-driven post-training combined with DMD technology increases inference speed by about 7.5 times, turning half a day's wait into instant output.
-
Conversational interactive creationDirector Agent transforms static generation into dynamic collaboration, supporting natural language planning, review, and local revisions, significantly lowering the barrier to creation.
-
High-definition real-time outputThe lightweight super-resolution module stably outputs high-resolution video with streaming latency, meeting the needs of professional content production.
-
Fully open sourceThe code and weights are all open source, built on LTX-2.3 and Gemma, and support academic research and secondary development.
JoyAI-Echo project address
- Project official website: https://echo-team-joy-future-academy-jd.github.io/Echo-LongVideo-Page/
- GitHub repository: https://github.com/jd-opensource/JoyAI-Echo
Comparison of JoyAI-Echo with similar competing products
| Comparison Dimensions | JoyAI-Echo | HappyOyster |
|---|---|---|
| Long video generation capability | Support the longest 5 minutesMulti-camera coherent story generation | It supports the generation of long videos, but the specific duration is not explicitly disclosed. |
| Role/Identity Consistency | 59.4% User preferences; cross-modal memory ensures consistency between the appearance and voice of characters in multiple camera shots. | 27.7% of users prefer this method; similar memory mechanisms are not explicitly disclosed. |
| Visual Aesthetics | 63.6% User preferences | 27.6% of users' preferences |
| Audio quality | 81.7% User preferences; joint audio and video generation, stable sound quality. | 11.8% of user preferences |
| Prompt words follow | 80.6% User preferences; Director Agent automatically splits scripts and shots. | 5.9% of users' preferences |
| Generation speed | DMD distillation is accelerated.7.5 timesInference speedup, supports streaming generation | Standard multi-step diffusion inference does not explicitly disclose the acceleration mechanism. |
| Conversational editing | Director Agent supports natural language interaction and local shot revisions, eliminating the need to rewatch the entire film. | Dialogic local editing is not explicitly supported. |
| Real-time super-resolution | Lightweight single-step super-resolution, supporting up to 1472×2560 | Real-time super-resolution is not explicitly supported. |
| Open source situation | The code and weights are fully open source (for academic research/non-commercial use). | Not open source |
| Underlying architecture | Based on LTX-2.3 + Gemma-3-12B, conditional generation of paired cross-modal memory banks. | Based on a self-developed model, few specific technical details have been disclosed. |
Application scenarios of JoyAI-Echo
-
Virtual story creation and animation productionIt generates coherent animated stories lasting several minutes, maintaining a high degree of consistency in character appearance, voice, and personality across multiple shots, significantly reducing the cost of traditional animation production.
-
Digital Human Content Production and Live StreamingIt enables the rapid generation of long video content for virtual anchors and digital human customer service representatives, ensuring that the digital human's face and voice do not drift during long-term output, thereby enhancing realism and professionalism.
-
Brand marketing videos iterate rapidlyWith Director Agent's conversational editing capabilities, marketing teams can modify ad scripts and shots as easily as chatting, quickly producing multiple versions of brand videos and shortening the creative cycle.
-
Film and television pre-production demonstration and storyboard productionDirectors and producers can use natural language to generate storyboards and preview videos for feature films, verifying camera language, character movement, and narrative rhythm before formal shooting, thus reducing trial and error costs.