01 / Overview
Video fusion aims to integrate complementary information from multiple source videos while preserving temporal consistency. Effective modeling of temporal dynamics is essential to this goal, yet existing methods incur substantial computational overhead from optical flow estimation and feature warping. In this paper, we present MambaVF, an efficient video fusion framework that uses state space model (SSM) to achieve temporal modeling without explicit motion estimation. First, by formulating video fusion as a sequential state update process, MambaVF captures long-range temporal dependencies with linear complexity, significantly reducing computation and memory costs. Second, the lightweight SSM-based fusion module eliminates conventional flow-guided alignment. Instead, it introduces a mutual state fusion module and a spatio-temporal bidirectional scanning mechanism to enable information aggregation across video streams. Experiments on multiple benchmarks confirm that MambaVF reaches state-of-the-art performance in different video fusion applications (multi-exposure, multi-focus, infrared-visible, medical), while reducing parameters by >90% and FLOPs by >80%, resulting in >50% shorter runtime.
03 / Architecture
Detailed illustration of our MambaVF architecture.
Exchanges complementary information between source streams through reciprocal state interaction, without optical flow estimation or feature warping.
Captures spatial details and temporal dependencies through eight bidirectional scanning paths shared by both sources.
04 / Visualization
05 / Benchmarks
Quantitative evaluation results for the Multi-Exposure Fusion (540p) and Multi-Focus Fusion (480p) task. The red and blue highlights indicate the highest and second-highest scores
Quantitative evaluation results for the Infrared-Visible Fusion and Medical Video Fusion task. The red and blue highlights indicate the highest and second-highest scores
Refer to the main paper linked above for more details on qualitative, quantitative, and ablation studies.
06 / A closer look
Multi-Exposure Video Fusion Comparison
07 / Citation
@article{zhao2026mambavfstatespacemodel,
title={MambaVF: State Space Model for Efficient Video Fusion},
author={Zixiang Zhao and Yukun Cui and Lilun Deng and Haowen Bai and Haotong Qin and Tao Feng and Konrad Schindler},
journal={arXiv preprint arXiv:2602.06017},
year={2026},
}