News
DeepSeek, in collaboration with Peking University, has open-sourced DSpark, a framework for accelerating speculative decoding.
DeepSeek, in collaboration with Peking University, has launched DSpark, a framework for accelerating large model inference, which has been integrated into the DeepSeek-V4 series production systems. With the same total throughput, DeepSeek-V4-Flash offers a 60%–85% speedup for single-user generation, while DeepSeek-V4-Pro achieves a 57%–78% speedup. DSpark employs a semi-autoregressive architecture and a confidence-based scheduling verification mechanism, balancing draft generation speed and coherence, and dynamically adjusting the verification length based on system load.