Python Apache-2.0

Flash-VStream

This is the official implementation of ICCV 2025 "Flash-VStream: Efficient Real-Time Understanding for Long Video Streams"

I

IVGSZ

Dernière activité 15 oct. 2025
IVGSZ/Flash-VStream

288

étoiles

22

forks

9

issues ouvertes

Ce README est souvent en anglais.

Flash-VStream Logo

[ICCV 2025] Flash-VStream: Efficient Real-Time Understanding for Long Video Streams

Haoji Zhang*, Yiqin Wang*, Yansong Tang✉, Yong Liu, Jiashi Feng, Xiaojie Jin✉†

*Equally contributing first authors, ✉Correspondence, †Project Leader

Work done when interning at Bytedance.

We proposed Flash-VStream, an efficient VLM with a novel Flash Memory mechanism that enables real-time understanding and Q&A of extremely long video streams. Our model achieves outstanding accuracy and efficiency on EgoSchema, MLVU, LVBench, MVBench and Video-MME Benchmarks.

News

Contents

Flash-VStream-Qwen

See Flash-VStream-Qwen/README.md.

Flash-VStream-LLaVA

See Flash-VStream-LLaVA/README.md.

Citation

If you find this project useful in your research, please consider citing:

@article{zhang2025flashvstream,
    title={Flash-VStream: Efficient Real-Time Understanding for Long Video Streams}, 
    author={Haoji Zhang and Yiqin Wang and Yansong Tang and Yong Liu and Jiashi Feng and Xiaojie Jin},
    journal={arXiv preprint arXiv:2506.23825},
    year={2025},
}
@article{zhang2024flashvstream,
    title={Flash-vstream: Memory-based real-time understanding for long video streams},
    author={Zhang, Haoji and Wang, Yiqin and Tang, Yansong and Liu, Yong and Feng, Jiashi and Dai, Jifeng and Jin, Xiaojie},
    journal={arXiv preprint arXiv:2406.08085},
    year={2024}
}

Acknowledgement

We would like to thank the following repos for their great work:

  • This work is built upon the LLaVA.
  • This work utilizes LLMs from Vicuna.
  • Some code is borrowed from LLaMA-VID.
  • We perform video-based evaluation from Video-ChatGPT.

License

Code License

This project is licensed under the Apache-2.0 License.

Projets similaires

StreamingVLM: Real-Time Understanding for Infinite Video Streams

Python
Mmit-han-lab
1,1 k étoiles69

Official Pytorch implementation of StreamV2V.

Python
JJeff-LiangF
549 étoiles59

PySlowFast: video understanding codebase from FAIR for reproducing state-of-the-art video models.

Python
Ffacebookresearch
7,4 k étoiles1,3 k