Planning + reasoning challenge by Motional and UCLA
VideoLLaMA2-AV
Unified MLLM with Text-Aligned Representations