SlowFast-LLaVA-1.5 - Apple's multimodal long video understanding model
SlowFast-LLaVA-1.5 (SF-LLaVA-1.5 for short) is a high-efficiency video language model designed specifically for long video understanding. Based on a two-stream (SlowFast) mechanism, it balances processing more input frames with reducing the number of tokens per frame...