Keye-VL-2.0-30B-A3B - Kuaishou's open-source self-developed multimodal large model
Keye-VL-2.0-30B-A3B is a self-developed multimodal large model open-sourced by Kuaishou, serving as the main foundation for 30B-level models. This model is the first to introduce DSA sparse attention into multimodal scenarios, supporting ultra-long contexts of up to 256K, enabling hour-level video...
What is Keye-VL-2.0-30B-A3B?
Keye-VL-2.0-30B-A3B is a self-developed multimodal large model open-sourced by Kuaishou, serving as the main foundation for 30B-level models. This model is the first to introduce DSA sparse attention into multimodal scenarios, supporting ultra-long contexts of up to 256K, enabling millisecond-level temporal inference for hour-level videos. It surpasses Gemini-2.5-Pro and Gemini 3 Flash in TimeLens benchmark tests and unlocks Agent collaboration mechanisms such as Code, Tool, and Search for the first time, allowing the model to evolve from an observer to an actor.
Main functions of Keye-VL-2.0-30B-A3B
-
Understanding Long VideosIt supports 256K ultra-long contexts, can process hour-long video sequences, and achieves near-lossless deep temporal inference.
-
Temporal causal reasoning: Capture the causal chain behind the scene in a continuous temporal flow, and realize the leap from "seeing the scene" to "understanding the logic".
-
Millisecond-level frame-level positioningIt possesses surgical-level fine-grained resolution capabilities, enabling the breakdown of complex processes or game highlights with precision down to the timestamp.
-
Cross-modal deep fusionIt simultaneously processes visual, audio, and textual information to achieve collaborative understanding and deep semantic alignment across multiple modalities.
-
Agent Collaborative ExecutionFor the first time, it unlocks system-level autonomous collaboration and task execution capabilities in complex scenarios such as code generation, tool invocation, and search.
-
High-noise information purificationIt accurately captures keyframes and clarifies dynamic patterns in complex scenes, effectively filtering redundant information and retaining core content.
Technical Principles of Keye-VL-2.0-30B-A3B
- DSA Sparse Attention MechanismThis is the first time that DeepSeek Sparse Attention has been introduced into multimodal understanding, and the combination of sparse attention and targeted feature aggregation has broken through the exponential computing power bottleneck of ultra-long visual contexts.
- Ultra-long context architectureIt adopts a 256K token-level end-to-end architecture to achieve coherent depth perception of long video sequences without segmentation or truncation.
- Fine-grained timing understanding engineThrough frame-level action boundary recognition, dynamic visual analysis, and audio-visual collaborative modeling, millisecond-level accurate temporal positioning and causal inference are achieved.
- Agent Collaboration FrameworkIt integrates Code Interpreter, Tool Use, and Search capabilities to build a closed-loop decision-making system that goes from multimodal perception to logical reasoning and then to tool execution.
- Unified Multimodal Feature FusionIt maps visual, audio, and text features to a shared representation space, enabling deep semantic alignment and joint reasoning of cross-modal information.
How to use Keye-VL-2.0-30B-A3B
-
Get the modelThe fully open-source model weights and deployment documentation can be downloaded from GitHub, Hugging Face, or ModelScope.
-
Hardware preparationIt requires an H800 or equivalent graphics card and uses at least two GPUs for multi-GPU tensor parallel inference.
-
Docker rapid deploymentYou can directly pull the official Docker image and run it to complete the environment configuration and model loading with one click.
-
Source code installation and deploymentClone the three dependency repositories in sequence: Keye customized version SGLang, DeepGEMM, and EffectiveKernels, and then compile and install them.
-
Start the inference serviceBy using SGLang to load model weights, setting tensor parallel parameters, and enabling remote code trust, you can start an API service compatible with the OpenAI protocol locally.
-
Call APIAfter startup, video and text instructions are sent via standard HTTP requests. The model will return structured long video understanding results or Agent execution output.
Keye-VL-2.0-30B-A3B's core advantages
-
DSA first implementation in multimodalThis is the first time that DeepSeek Sparse Attention has been introduced into a multimodal understanding scenario, fundamentally breaking through the exponential computing power bottleneck caused by ultra-long visual contexts and enabling efficient inference of hour-long videos.
-
256K Extremely Long ContextIt supports token-level ultra-long contexts of up to 256K, enabling near-lossless end-to-end depth perception of hourly video sequences without the need for segmentation and truncation as in traditional models.
-
Millisecond-level frame-level positioningIt possesses surgical-level fine-grained timing analysis capabilities, enabling precise breakdown and positioning of every key action in complex processes, game highlights, and other scenarios, down to the time stamp.
-
Temporal causal reasoningBeyond simple image label recognition, it captures causal chains in continuous temporal flow, achieving a leap from "seeing the image" to "understanding the logic." For example, it can directly infer the safety strategy of "group tours are better than self-driving" from images of "car accidents in snowy conditions."
-
Agent collaboration mechanismThe Keye series unlocks system-level autonomous collaboration and execution capabilities for complex scenarios such as Code, Tool, and Search for the first time, allowing models to evolve from passive "observers" to proactive "actors" that solve tasks.
Project address for Keye-VL-2.0-30B-A3B
- GitHub repositoryhttps://github.com/Kwai-Keye/Keye
- HuggingFace model libraryhttps://huggingface.co/Kwai-Keye/Keye-VL-2.0-30B-A3B
Comparison of Keye-VL-2.0-30B-A3B with similar competing products
| Comparison Dimensions | Keye-VL-2.0-30B-A3B | Gemini-2.5-Pro | Gemini 3 Flash |
|---|---|---|---|
| Company | Kuaishou | ||
| Model size | 30B | Not yet released (Pro level) | Unreleased (Flash-level) |
| Core Architecture | DSA Sparse Attention + Multimodal Fusion | Closed-source multimodal architecture | Closed-source multimodal architecture |
| Extremely long context | 256K Token(Hour-long video) | Long context | Long context |
| ActivityNet-TimeLens< Video motion positioning |
mIoU 58.5 | mIoU 58.1 | mIoU 57.0 |
| Charades-TimeLens< Analysis of the timing of daily actions |
mIoU 58.4 | — | mIoU 61.2 |
| QVHighlights-TimeLens< Highlight extraction |
mIoU 70.1 | — | mIoU 49.5 |
| Agent Collaboration Capability | First unlock< Code / Tool / Search |
support | support |
| Open source situation | Fully open source< (Weight + Code + Documentation) |
Closed source | Closed source |
Application scenarios of Keye-VL-2.0-30B-A3B
-
Long video content comprehensionKeye-VL-2.0-30B-A3B can perform in-depth temporal causal reasoning on hour-long videos such as travel vlogs, documentaries, and instructional videos, and automatically generate a complete structured summary that includes equipment suggestions, budget planning, attraction recommendations, and safety tips.
-
Industrial process analysisThis model can locate key action nodes in complex process videos with millisecond-level accuracy, accurately break down the manufacturing process into multiple stages and mark them with timestamps, and is suitable for process decomposition, operation specification extraction and quality inspection process optimization.
-
esports and sports content productionBased on a deep understanding of visual tension, audio-visual synergy, and narrative logic, the model can accurately identify highlight moments and emotional resonance points in e-sports or sports event videos, achieving intelligent extraction of exciting moments that goes beyond simple kill prompts.
-
Agent Automated TasksAs the first collaborative mechanism unlocked in the Keye series, this model supports system-level autonomous execution of code generation, tool invocation, and multi-step search, enabling it to complete complex closed-loop tasks from multimodal perception to logical reasoning and then to tool invocation.
-
Education and TrainingIn practical teaching scenarios, the model can locate key actions and break down steps in student operation videos at the millisecond level, providing teachers with accurate teaching feedback and operational correction basis, and assisting in skills assessment and course optimization.