Skip to main navigation Skip to search Skip to main content

LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

  • Xiaoqian Shen
  • , Yunyang Xiong
  • , Changsheng Zhao
  • , Lemeng Wu
  • , Jun Chen
  • , Chenchen Zhu
  • , Zechun Liu
  • , Fanyi Xiao
  • , Balakrishnan Varadarajan
  • , Florian Bordes
  • , Zhuang Liu
  • , Hu Xu
  • , Hyunwoo J. Kim
  • , Bilge Soran
  • , Raghuraman Krishnamoorthi
  • , Mohamed Elhoseiny
  • , Vikas Chandra

Research output: Contribution to journalConference articlepeer-review

Abstract

Multimodal Large Language Models (MLLMs) have shown promising progress in understanding and analyzing video content. However, processing long videos remains a significant challenge constrained by LLM’s context size. To address this limitation, we propose LongVU, a spatiotemporal adaptive compression mechanism that reduces the number of video tokens while preserving visual details of long videos. Our idea is based on leveraging cross-modal query and inter-frame dependencies to adaptively reduce temporal and spatial redundancy in videos. Specifically, we leverage DINOv2 features to remove redundant frames that exhibit high similarity. Then we utilize text-guided cross-modal query for selective frame feature reduction. Further, we perform spatial token reduction across frames based on their temporal dependencies. Our adaptive compression strategy effectively processes a large number of frames with little visual information loss within given context length. Our LongVU consistently surpass existing methods across a variety of video understanding benchmarks, especially on hour-long video understanding tasks such as VideoMME and MLVU. Given a lightweight LLM, our LongVU also scales effectively into a smaller size with state-of-the-art video understanding performance.

Original languageEnglish (US)
Pages (from-to)54582-54599
Number of pages18
JournalProceedings of Machine Learning Research
Volume267
StatePublished - 2025
Externally publishedYes
Event42nd International Conference on Machine Learning, ICML 2025 - Vancouver, Canada
Duration: Jul 13 2025Jul 19 2025

All Science Journal Classification (ASJC) codes

  • Software
  • Control and Systems Engineering
  • Statistics and Probability
  • Artificial Intelligence

Fingerprint

Dive into the research topics of 'LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding'. Together they form a unique fingerprint.

Cite this