AB
AiBoss
News

ByteDance's video model Vidi2 surpasses Gemini 3 Pro! Its comprehension capabilities are off the charts.

ByteDance has released its next-generation video understanding model, Vidi2, which outperforms GPT-5 and Gemini 3 Pro in core tasks such as spatiotemporal localization. The model can accurately understand hours of long video content and directly generate complete JSON editing solutions that include details such as editing time points, subtitles, and background music, enabling AI-automated editing from raw footage to finished product.