AB
AiBoss
News

Inco AI releases DFlash 2, improving the efficiency of parallel drafting in speculative decoding.

Inco AI has released DFlash 2. According to their tests, with an increase of about 1% in single-cycle latency, the acceptable output per validation cycle is improved by more than 20%, and the improvement on different benchmarks is about 16% to 25%.

Inco AI releases DFlash 2 to improve the efficiency of parallel drafting in speculative decoding.

According to the company's published tests, with an increase of approximately 1% in single-cycle latency, the acceptable output per validation cycle increased by more than 20%, with improvements of approximately 16% to 25% on different benchmarks. These figures are from publisher tests, and actual gains will be affected by the model, hardware, and workload.

refer to:Inco AI Official Announcement