🐍 MambaVF 🐍

MambaVF: State Space Model for Efficient Video Fusion

1 ETH Zürich  2 Xi'an Jiaotong University  3 Nanyang Technological University
4 Hong Kong Polytechnic University  5 Tsinghua University
5.7%Parameters vs. UniVF
13.4%FLOPs vs. UniVF
2.1×Inference speedup

Compared with UniVF (NeurIPS'25), our MambaVF not only attains state-of-the-art performance on VF-Bench, but requires only 5.7% of the parameters and 13.4% of the FLOPs, while achieving a 2.1× speedup.

01 / Overview

Abstract

Video fusion aims to integrate complementary information from multiple source videos while preserving temporal consistency. Effective modeling of temporal dynamics is essential to this goal, yet existing methods incur substantial computational overhead from optical flow estimation and feature warping. In this paper, we present MambaVF, an efficient video fusion framework that uses state space model (SSM) to achieve temporal modeling without explicit motion estimation. First, by formulating video fusion as a sequential state update process, MambaVF captures long-range temporal dependencies with linear complexity, significantly reducing computation and memory costs. Second, the lightweight SSM-based fusion module eliminates conventional flow-guided alignment. Instead, it introduces a mutual state fusion module and a spatio-temporal bidirectional scanning mechanism to enable information aggregation across video streams. Experiments on multiple benchmarks confirm that MambaVF reaches state-of-the-art performance in different video fusion applications (multi-exposure, multi-focus, infrared-visible, medical), while reducing parameters by >90% and FLOPs by >80%, resulting in >50% shorter runtime.


03 / Architecture

How MambaVF Works

Detailed illustration of our MambaVF architecture.

Mutual State Fusion (MuFuse) Mamba Encoder

Exchanges complementary information between source streams through reciprocal state interaction, without optical flow estimation or feature warping.

Spatio-Temporal Bidirectional (STB) Scanning

Captures spatial details and temporal dependencies through eight bidirectional scanning paths shared by both sources.

04 / Visualization

Qualitative Video Comparison


05 / Benchmarks

Quantitative comparison with other methods

Quantitative evaluation results for the Multi-Exposure Fusion (540p) and Multi-Focus Fusion (480p) task. The red and blue highlights indicate the highest and second-highest scores

Quantitative evaluation results for the Infrared-Visible Fusion and Medical Video Fusion task. The red and blue highlights indicate the highest and second-highest scores

Refer to the main paper linked above for more details on qualitative, quantitative, and ablation studies.

06 / A closer look

Qualitative Image Comparison

Multi-Exposure Video Fusion Comparison

07 / Citation

Citation


  @article{zhao2026mambavfstatespacemodel,
    title={MambaVF: State Space Model for Efficient Video Fusion},
    author={Zixiang Zhao and Yukun Cui and Lilun Deng and Haowen Bai and Haotong Qin and Tao Feng and Konrad Schindler},
    journal={arXiv preprint arXiv:2602.06017},
    year={2026},
  }