Keling O1 - Keling AI's first unified multimodal video generation model
Keling O1 (Keling Video O1 Model) is the world's first unified multimodal video generation model launched by Keling AI. The model achieves seamless integration of video generation, editing and understanding through an innovative multimodal visual language (MVL) architecture.
What is Keling O1?
Keling O1 (Keling Video O1 Model) is the world's first unified multimodal video generation model launched by Keling AI. Through an innovative multimodal visual language (MVL) architecture, the model achieves seamless integration of video generation, editing, and understanding. Supporting multimodal inputs such as images, videos, and text, the model enables versatile creative editing, solves the challenge of video consistency, and provides a variety of creative combinations. Users can generate accurate video content through simple dialogue, exploring limitless creative possibilities.
The latest upgrade to the Keling O1 model adds a 720p mode and supports 3-10 second free narrative, bringing more control and freedom to creation.
Main functions of Keling O1
-
All-round engineKeling O1 is the world's first unified multimodal video model, which can complete the entire creation process of video generation, editing and modification in one stop without switching between multiple tools.
-
Omnipotent CommandIt supports multimodal input, including images, videos, and text. Through deep semantic understanding, users can easily generate and edit video content through simple conversations.
-
All-round referenceBy constructing subjects from multiple perspectives and freely combining multiple subjects, the problem of video consistency is solved, ensuring that the video footage remains accurate and coherent regardless of how the camera moves.
-
Super CombinationIt supports the combined use of different skills, such as adding a subject and modifying the background at the same time, generating multiple creative variations at once, and exploring endless creative possibilities.
-
Control the rhythmIt supports video lengths of 3-10 seconds, allowing users to freely control the video's rhythm.
-
Added 720p modeWhile retaining the original 1080p core capabilities, a new 720p mode has been added, which is suitable for lightweight creation and reduces equipment requirements.
-
Free narrative durationThe first and last frames support free narration of 3-10 seconds, breaking the fixed length limit. Creators can freely define the beginning and end length of the video, improving creative flexibility.
The technical principle of Keling O1
-
New video generation modelBreaking away from the functional fragmentation of traditional video models, a new generative foundation is constructed, integrating Multimodal Transformer and Multimodal Long Context for multimodal understanding.
-
Multimodal Visual Language (MVL)Introducing MVL as an interaction medium, it achieves deep integration of text semantics and multimodal signals through Transformer, supporting flexible and seamless integration of multiple tasks within a single input box.
-
Intelligent reasoning abilityBased on MVL input, the model achieves accurate multimodal reference and highly flexible interactive editing, supporting long contexts and temporal narratives. Combined with Chain-of-thought technology, the model possesses common-sense reasoning and event deduction capabilities, demonstrating intelligent performance in video generation.
Performance of Keling O1
-
Image reference task:In the image reference task, the model achieved an overall win rate of 247%, indicating excellent performance in both overall results and multiple sub-dimensions.Compared to Google Veo 3.1's Ingredients to Video, the Video O1 model significantly outperforms the image reference task.
-
Instruction translation task:In the instruction transformation task, the model achieved an overall win-loss ratio of 230%, demonstrating excellent performance in both overall results and multiple sub-dimensions.Compared to Runway Alph, the model also significantly outperforms it in instruction transformation tasks.
How to use Keling O1
- Access PlatformVisit the Keling official website or Keling App to complete user account registration and login.
- Select ModelSelect the video O1 model on the platform.
- Upload materialsUpload reference images, video clips, text descriptions, and other materials as needed.
- Input commandUse the multimodal instruction input area to input creation instructions.
- Generate videoThe model generates a video based on the provided materials and instructions. The video length can be specified, such as 3-10 seconds.
- Editing and adjustingUse the editing functions provided by the model, such as adding, deleting, and modifying video content, and switching shot types/perspectives.
- Preview and ExportPreview the generated video to ensure it meets your requirements. Once satisfied, export the video to your local device.
Application scenarios of Keling O1
- Social media content creationUsers can quickly generate short videos suitable for social media platforms such as TikTok and Instagram for personal sharing or brand marketing.
- Online education and trainingEducators can create interactive video courses and training materials to enhance the appeal and effectiveness of distance learning.
- Advertising and marketing videosBusinesses and marketing teams use models to generate engaging advertising videos for product promotion and brand building.
- Film and video productionFilmmakers and video editors use models for pre-production, such as creating storyboards, proof-of-concept, and animation effects.
- Corporate promotion and demonstrationCompanies produce high-quality promotional videos and demonstration videos for use in company introductions, product demonstrations, and event reporting, thereby enhancing their corporate image.