News
Inco AI releases DFlash 2, improving the efficiency of parallel drafting in speculative decoding.
Inco AI has released DFlash 2. According to their tests, with an increase of about 1% in single-cycle latency, the acceptable output per validation cycle is improved by more than 20%, and the improvement on different benchmarks is about 16% to 25%.
Inco AI releases DFlash 2 to improve the efficiency of parallel drafting in speculative decoding.
According to the company's published tests, with an increase of approximately 1% in single-cycle latency, the acceptable output per validation cycle increased by more than 20%, with improvements of approximately 16% to 25% on different benchmarks. These figures are from publisher tests, and actual gains will be affected by the model, hardware, and workload.
refer to:Inco AI Official Announcement