Light Mode

MAC-I2: Learned Metrics-Aware Covariance for Robust Visual-Inertial Fusion in Initialization and Calibration

Xiang Fei*1,  Yuheng Qiu*1,   Can Xu1,3,  Yutian Chen1,   Ruogu Li1,   Xingxing Zuo2,  Wenshan Wang1,  Sebastian Scherer1

* Equal Contribution

1 Robotics Institute, Carnegie Mellon University

2 MBZUAI

3 UTIAS, University of Toronto

Background: trajectories estimated during calibration

Ground TruthMAC-I2 Estimated

Abstract

Visual-Inertial (VI) fusion is fundamental to accurate and robust state estimation, where camera and IMU measurements are combined according to their respective uncertainties. Existing methods, however, fuse the two modalities with predefined uncertainties, regardless of how reliable each is in the local context, and thus often struggle under challenging environments involving illumination changes, dynamic objects, and textureless regions. In this paper, we present MAC-I2, which achieves robust VI fusion through learned metrics-aware covariance for both modalities, so that vision and IMU compete on their own merits rather than relying on predefined uncertainties. Here, metrics-aware means that each predicted covariance faithfully reflects the actual magnitude of the corresponding measurement noise. On the visual side, we propagate learned feature-matching uncertainties into pose covariances for the fusion. On the inertial side, motivated by the observation that integration error accumulates sharply at the early stage and grows slowly afterward, we design a learned IMU model with a learnable initial covariance, and propose a dedicated fine-tuning strategy on a held-out training subset to enable the metrics-aware covariance on unseen sequences. As a showcase, we build a VI initialization and calibration system, since accurate and robust initialization and calibration are the prerequisite for any reliable VI system. Experiments on EuRoC and VBR show that MAC-I2 substantially outperforms existing methods: it achieves a 99.9% initialization success rate on EuRoC, reducing gravity and velocity errors by about 60% and 42% over the strongest baseline, and maintains an 80% success rate on challenging VBR sequences where baseline methods such as VINS-Mono drop below 10%.

99.9%Initialization success rate on EuRoC
↓ 60%Gravity error vs. strongest baseline
↓ 42%Velocity error vs. strongest baseline
80%Success rate on challenging VBR sequences

Demonstrations

MAC-I2 stays reliable where predefined uncertainties fail. Explore its behavior across challenging environments and real-world deployments.

Illumination Change Extreme Exposure

Dynamic Scene Moving Objects

Dark Room Dark

Bottom-left shows the input images.

Calibration

For each sequence we show the sensor input (left) and the estimated trajectory during calibration (right), played in sync.

Sequence 1

Sensor Input

Estimated Trajectory during Calibration

Ground TruthMAC-I2 Estimated

Sequence 2

Sensor Input

Estimated Trajectory during Calibration

Ground TruthMAC-I2 Estimated

Sequence 3

Sensor Input

Estimated Trajectory during Calibration

Ground TruthMAC-I2 Estimated

Sequence 4

Sensor Input

Estimated Trajectory during Calibration

Ground TruthMAC-I2 Estimated

Method

MAC-I2 learns metrics-aware covariance for both modalities, so that each measurement is weighted by its actual reliability in the local context rather than by a predefined rule.

Metrics-Aware Visual Covariance

On the visual side, we propagate the learned feature-matching uncertainties from MAC-VO into pose covariances through the information matrix at the convergence of the visual pose estimation, bringing metrics-aware visual uncertainty into the fusion.

Color Inversion ON
Figure 1. The learned visual pose covariance tracks the actual magnitude of the pose error across the sequence.

Metrics-Aware Inertial Covariance

On the inertial side, the integration error accumulates sharply at the early stage of the integration window and grows slowly afterward. We design a learned IMU model with a learnable initial covariance to capture this pattern, together with a held-out fine-tuning strategy that keeps the predicted covariance metrics-aware on unseen sequences.

Color Inversion ON
Figure 2. The learned inertial covariance faithfully reflects the actual magnitude of the IMU integration error.

Quantitative Results

Color Inversion ON
Figure 3. Gravity error vs. initialization success rate on EuRoC — MAC-I2 reaches a 99.9% success rate with the lowest gravity error.
Color Inversion ON
Figure 4. Calibration accuracy across sequences compared with existing methods.

Citation

@misc{fei2026maci2learnedmetricsawarecovariance,
    title={MAC-I$^2$: Learned Metrics-Aware Covariance for Robust Visual-Inertial Fusion in Initialization and Calibration},
    author={Xiang Fei and Yuheng Qiu and Can Xu and Yutian Chen and Ruogu Li and Xingxing Zuo and Wenshan Wang and Sebastian Scherer},
    year={2026},
    eprint={2609.07116},
    archivePrefix={arXiv},
    primaryClass={cs.RO},
    url={https://arxiv.org/abs/2609.07116},
  }