SIGGRAPH-ASIA 2026

CompMVR:
Compression-Aware Multi-View Restoration Using Diffusion Models for Geometrically Consistent 3D Reconstruction

Dong-hwi Kim*1, Chaewon Moon*1, Hojun Song1, Dongbeom Kim1, Junyeong Jang1, Aro Kim1, Gahyeon Kim1, Heejung Choi1, Jehee Kim1, Gianella Cravioto1, Sohyun Lee1, Gyeongjin Choi1, EunHye Jeong1, Soo Ye Kim†2, Jaehyup Lee†1, Sang-hyo Park†1

1Kyungpook National University, South Korea    |    2Adobe Research, USA

* Equal contribution    |    † Corresponding author

Problem overview

Abstract

Real-world 3D reconstruction pipelines often rely on multi-view images degraded by lossy compression. While compression artifacts have been extensively studied in single-image or video restoration, their impact on multi-view geometry remains largely overlooked. In this work, we identify compression-induced cross-view geometric inconsistency as a critical failure mode, where view-dependent artifacts disrupt feature matching, dense correspondence, camera parameter estimation, and downstream 3D reconstruction tasks such as novel view synthesis.

To address this problem, we propose CompMVR, a compression-aware multi-view diffusion restoration framework and compressed datasets that jointly refines compressed views while preserving cross-view consistency. Unlike conventional restoration methods that optimize each image independently for 2D perceptual quality, CompMVR leverages learned compression priors and cross-view correspondence to recover visual details while improving geometric reliability across views.

Experiments across diverse datasets, codecs, and compression levels demonstrate gains in restoration quality and downstream 3D tasks, including novel-view synthesis, camera pose estimation, and view matching. These results establish compressed multi-view restoration as a distinct problem and highlight its importance for 3D reconstruction under lossy compression.

Method

Overview of the CompMVR framework

(a) In Stage 1, we learn a compression-aware latent representation by training a compression prior embedder. The latent is optimized through image reconstruction and supervision of coding parameters (codec type and QP). (b) In Stage 2, we perform multi-view restoration using a diffusion model conditioned on the learned compression prior. The latent is injected via cross-attention and fused with the UNet input, while a compression artifact estimator (CAE) predicts spatially varying residual maps for degradation-aware restoration. To enforce cross-view geometric consistency, we further apply a cross-view correspondence loss during training.

Experiments

Quantitative Results

2D Restoration

2D restoration and cross-view consistency results
CompMVR achieves competitive 2D restoration quality and the best cross-view consistency, demonstrating superior geometric reliability over state-of-the-art single-image and video restoration methods.

3D Downstream Tasks

Novel-view synthesis results
NVS (3DGS). Across diverse codec and QP settings, CompMVR consistently achieves the best novel-view synthesis performance among restoration methods, with higher PSNR/SSIM and lower LPIPS, demonstrating improved downstream 3D reconstruction quality.
VGGT camera pose estimation results
Camera Pose Estimation. CompMVR consistently achieves the best camera pose estimation accuracy across diverse codec and QP settings, demonstrating that our restoration better preserves geometrically reliable cross-view information.
View matching results
View Matching. CompMVR improves view matching across compressed views, demonstrating that restoring cross-view consistency leads to more reliable correspondences for downstream geometric tasks.

Qualitative Results

2D Restoration

2D_results
2D restoration qualitative comparison on the Free dataset. The top rows correspond to an indoor scene (Lab), and the bottom rows correspond to an outdoor scene (Sky). The Lab scene is compressed using HEVC with QP 42, while the Sky scene is compressed using AV1 with QP 42. Highlighted regions reveal that our method better preserves fine details and structures under severe compression

3D Novel-View Synthesis

3DGS_render
3D novel-view synthesis results rendered at ground-truth target camera poses. The top two rows correspond to HEVC QP42, and the bottom two rows correspond to AVC QP42. We compare each method with the corresponding GT view using full-scene renderings and zoomed-in regions. Our method better preserves object structures, text regions, and high-frequency details, producing results closer to GT than diffusion-based restoration baselines. Red boxes indicate representative regions for comparison.
3DGS_viewer
3D novel-view synthesis results from restored multi-view inputs on the Free dataset. We evaluate global consistency using scene renderings synthesized from multiple viewpoints and local structural fidelity using zoomed-in regions. Views 1 and 2 correspond to HEVC QP42, and View 3 corresponds to AV1 QP42. Our method preserves sharper geometry-critical details and reduces blur and distortion compared with diffusion-based restoration baselines. Red boxes indicate representative regions for comparison.

BibTeX

          {kim2026compmvr,
            author = {Kim, Dong-hwi and others},
            title = {CompMVR: Compression-Aware Multi-View Restoration Using Diffusion Models 
                      for Geometrically Consistent 3D Reconstruction},
            booktitle = {SIGGRAPH Asia 2026 Conference Papers},
            year = {2026},
            address = {Kuala Lumpur, Malaysia},
            publisher = {Association for Computing Machinery},
            doi = {10.1145/3829340.3842283},
            isbn = {979-8-4007-2842-6}
          }