面向暗光环境的鲁棒 3D 感知 RGB-NIR 成像新方法

HuggingFace Daily Papers(社区热门论文)·2026-07-31 08:00·24天前
AI 导读

东京大学等机构提出一种无需干净 RGB 监督的 3D 感知神经隐式融合模型,可在暗光下将极噪 RGB 观测与近红外(NIR)线索在 3D 空间中隐式融合,恢复干净 RGB 图像。该方法引入 NIR 调制的位置编码和 Color Code MLP,解决噪声过拟合与 NIR-RGB 映射不适定问题,在合成与真实多视图数据集上优于现有方法,并能泛化至不同噪声水平。代码与数据将公开。

HuggingFace Daily Papers(社区热门论文)
49AI 编辑部评分,满分 100

面向暗光环境的鲁棒 3D 感知 RGB-NIR 成像新方法

2026-07-31 08:00· 24天前
AI 导读

东京大学等机构提出一种无需干净 RGB 监督的 3D 感知神经隐式融合模型,可在暗光下将极噪 RGB 观测与近红外(NIR)线索在 3D 空间中隐式融合,恢复干净 RGB 图像。该方法引入 NIR 调制的位置编码和 Color Code MLP,解决噪声过拟合与 NIR-RGB 映射不适定问题,在合成与真实多视图数据集上优于现有方法,并能泛化至不同噪声水平。代码与数据将公开。

Muyao Niu

The University of Tokyo

Tokyo

Japan

muyao.niu@gmail.com

Mingze Ma

Adelaide University

Adelaide

Australia

mingze.ma@adelaide.edu.au

Yifan Zhan

The University of Tokyo

Tokyo

Japan

zhan-yifan@g.ecc.u-tokyo.ac.jp

Qingtian Zhu

The University of Tokyo

Tokyo

Japan

qtzhu@g.ecc.u-tokyo.ac.jp

Zhihang Zhong

School of Artificial Intelligence (SAI)

Shanghai

zhongzhihang@sjtu.edu.cn

Wei Guo

The University of Tokyo

Tokyo

Japan

guowei@g.ecc.u-tokyo.ac.jp

Chang Wen Chen

Hong Kong Polytechnic University

Hong Kong SAR

changwen.chen@polyu.edu.hk

Yinqiang Zheng

The University of Tokyo

Tokyo

Japan

yqzheng@ai.u-tokyo.ac.jp

Abstract.

Robust low-light imaging remains challenging for the community. Recent studies have explored fusing Near-Infrared (NIR) with noisy RGB to achieve improved enhancement, yet most methods depend on carefully curated training data pairs, with limited robustness under different scenarios. This paper offers a new perspective for RGB-NIR low-light imaging by incorporating 3D-aware neural modeling. Without using clean RGB supervision, a powerful model can be optimized to implicitly fuse extremely noisy RGB observations with NIR cues in 3D space, effectively recovering clean RGB images. The proposed model obviates the requirement for clean RGB data collection, generalizes across different noise levels. Extensive evaluations on synthetic and real data demonstrate its superiority. Codes available: https://github.com/MyNiuuu/3DarkFusion

1. Introduction

Robust dark imaging remains a fundamental challenge in computer vision. Despite rapid advances in camera hardware, noise interference under extreme low-light conditions continues to degrade image quality. Conventional strategies, such as long-exposure and burst photography, are prone to motion blur (Niu et al., 2026) and visibility issues (Xiong et al., 2021; Niu et al., 2023b), limiting their practicality.

To mitigate these issues, numerous low-light enhancement methods have been proposed, ranging from traditional denoising operators (Buades et al., 2005; Dabov et al., 2007; Dong et al., 2012b; Jia et al., 2019; Xu et al., 2018; Dong et al., 2012a; Gu et al., 2014) to modern deep learning–based networks (Zamir et al., 2021; Anwar and Barnes, 2019; Chen et al., 2022; Wang et al., 2022; Zamir et al., 2022). Among these, several image- and video-based RGB–NIR low-light enhancement methods have recently attracted increasing attention (Jin et al., 2022; Sheng et al., 2023; Wang et al., 2024; Xu et al., 2024) due to the invisibility of near-infrared signals and their complementary structural cues in darkness. However, existing approaches still face critical limitations. Supervised models are typically trained on “noisy RGB–NIR–clean RGB” triplets from specific domains, resulting in poor cross-domain generalization and limited robustness under varying noise levels. This leads to a natural question: Can robust low-light enhancement be achieved using only NIR and noisy RGB observations, without relying on any form of clean RGB supervision?

Refer to caption
Figure 1. Top: Given extremely noisy multi-view RGB observations in darkness, the proposed method recovers the clean RGB scene using Near-Infrared. NO clean RGB is used for optimizing the model. Bottom: Visual comparisons against methods from various paradigms on two scenes respectively from the synthetic dataset (“ROVER”) and the real-captured dataset (“SAKANA”).

This paper investigates a novel perspective on this question by leveraging a multi-view, 3D-aware neural implicit model for efficient RGB–NIR dark imaging without relying on clean RGB supervision. Building upon the volume rendering framework and the analysis of several intuitive baseline architectures, a 3D-aware neural implicit fusion architecture is carefully redesigned to fully exploit the potential of RGB-NIR modalities. In addition, an NIR-modulated positional encoding mechanism is introduced to address the inherent limitations of conventional positional encoding under severe noise, effectively suppressing noise-induced overfitting. Finally, a Color Code MLP is presented to resolve the fundamentally ill-posed mapping between NIR and RGB through a learned color code distribution. Comprehensive evaluations are conducted on both synthetic and real-world multi-view datasets containing NIR images and noisy RGB observations. The proposed framework demonstrates strong performance under extreme low-light conditions compared with alternative approaches. Moreover, it generalizes across varying noise levels compared to existing methods, without using the clean RGB for supervision. Codes and data will be released to support future research. The key contributions of this work are: 1) A new 3D-aware fusion model is proposed for RGB-NIR dark imaging. Different from existing RGB-NIR models, it effectively fuses noisy RGB observation with NIR cues in 3D space without requiring clean RGB supervision. 2) Based on the characteristic of RGB-NIR modalities, two insightful components are introduced to unlock the complementary potentials, further improving performance. 3) Both synthetic and real-world datasets are contributed, demonstrating the superiority of the proposed model across various scenarios.

2. Related Work

Low-light photography. Early low-light imaging methods rely on handcrafted techniques (Dong et al., 2012b; Xu et al., 2018; Dong et al., 2012a), while recent methods (Zamir et al., 2021; Anwar and Barnes, 2019; Chen et al., 2022; Wang et al., 2022) achieve superior performance with deep learning. For example, DnCNN (Zamir et al., 2021) employs CNN for denoising, while transformer-based methods like Restormer (Zamir et al., 2022) further improve the quality.

Denoising with multi-view models. Since the introduction of Neural Radiance Fields (NeRFs) (Mildenhall et al., 2021; Niu et al., 2024; Zhan et al., 2024), several studies (Mildenhall et al., 2022; Pearl et al., 2022; Wang et al., 2023) have investigated their potential for image denoising. RawNeRF (Mildenhall et al., 2022) performs novel-view synthesis directly on RAW data. LLNeRF (Wang et al., 2023) explores denoising in the sRGB domain. Despite these advancements, leveraging NIR for robust multi-view low-light imaging remains unexplored.

Guided image denoising. Beyond approaches that rely solely on RGB information for denoising, several studies (Eisemann and Durand, 2004; Krishnan and Fergus, 2009; Petschnigg et al., 2004; He et al., 2012; Yan et al., 2013; Deng and Dragotti, 2020; Li et al., 2019; Oh et al., 2023; Xia et al., 2021; Xiong et al., 2021; Jin et al., 2022; Sheng et al., 2023; Wang et al., 2024; Zhang et al., 2025; Xu et al., 2024) have explored the use of additional modalities, such as burst flash (Eisemann and Durand, 2004; Krishnan and Fergus, 2009; Petschnigg et al., 2004; He et al., 2012; Oh et al., 2023; Xiong et al., 2021) and Near-Infrared (NIR) imaging (Yan et al., 2013; Jin et al., 2022; Sheng et al., 2023; Wang et al., 2024; Xu et al., 2024; Wan et al., 2022; Wang et al., 2025; Niu et al., 2023a). SANet (Sheng et al., 2023) estimates a clean structure map for RGB-NIR fusion. NAID (Xu et al., 2024) further proposes a Selective Fusion Module.

3. Approach

3.1. Preliminary: MLPs and Volume Rendering

[Uncaptioned image]
媒体内容 · 前往原文查看
Figure 2. Architecture illustrations of (a) parallel structure and (b) NIR-conditioned structure.
[Uncaptioned image]
媒体内容 · 前往原文查看
Figure 3. Visual comparison results. From left to right: NIR observation, (a) result of parallel structure, (b) result of NIR-conditioned structure, and normal RGB image.

Previous research (Mildenhall et al., 2022; Pearl et al., 2022; Wang et al., 2023) indicates that integrating Multi-Layer Perceptrons (MLPs) with volume rendering can aid in image denoising. This ability arises from two factors: 1) MLPs inherently favor smooth, low-frequency predictions, and 2) the volume-rendering process enforces 3D consistency by aggregating information across multiple viewpoints. To validate this behavior, a vanilla NeRF architecture from (Mildenhall et al., 2022) was trained using noisy RGB inputs. The results (Fig. 4) show that although the method reduces noise to some extent, it introduces significant artifacts. While the volume-rendering framework remains robust and crucial with the enforcement of multi-view consistency, the characteristics of MLPs need to be carefully handled to fully exploit the potentials of RGB–NIR signals, particularly when supervised by noisy RGB data. The following sections retain the standard volume-rendering formulation without modification, while introducing a completely redesigned MLP architecture specifically tailored to the complementary properties of RGB–NIR modalities.

3.2. Fusing NIR and RGB in 3D Space

Refer to caption
Figure 4. Results of using vanilla NeRF MLP for our model.

Intuitive parallel structure. A straightforward parallel architecture jointly optimizes two separate MLPs for RGB and NIR observations. Since the primary purpose of incorporating NIR images is to provide structural guidance, the NIR MLP is used to estimate the volume density , while a gradient-stopping strategy is applied to prevent degraded RGB observations from adversely affecting its learning process. The architectural design and corresponding qualitative results are presented in Fig. 2 (a) and Fig. 3 (a), respectively. The additional structural information from NIR substantially improves the reconstruction of overall scene geometry compared with the vanilla NeRF MLP (Fig. 4). However, the model still fails to recover detailed 2D texture information already present in the NIR images, such as the background texture and the floor in the “Lego” scene. This limitation arises because the RGB MLP does not directly exploit the predicted NIR values, which encode rich texture information crucial for high-fidelity reconstruction.

[Uncaptioned image]
媒体内容 · 前往原文查看
Figure 5. Illustration of the final architecture with the Color Code MLP.
[Uncaptioned image]
媒体内容 · 前往原文查看
Figure 6. Visual comparison results. From left to right: NIR observation, (a) result without Color Code MLP, (b) result with Color Code MLP, and normal RGB image.
Refer to caption
Figure 7. (a) Applying P.E. to and in the RGB MLP introduces checkerboard artifacts. (b) Applying P.E. to instead effectively removes these artifacts and enhances the reconstruction of structural details.

NIR-conditioned structure. Building on the previous observations, the architecture is improved by feeding the NIR output into the RGB MLP while ensuring that gradients from the RGB MLP do not propagate back to the NIR MLP. This modification enables the RGB MLP to leverage the structural guidance provided by NIR while maintaining the integrity of the NIR representation. The revised architecture and its corresponding qualitative results are shown in Fig. 2 (b) and Fig. 3 (b), respectively. This enhanced design substantially improves both 3D structural reconstruction and the recovery of fine 2D texture details. By implicitly fusing NIR predictions with RGB predictions in 3D space, the model effectively exploits additional structural information for improved results.

3.3. Revisiting Positional Encoding with NIR

Positional Encoding (P.E.) transforms 3D coordinates and 2D view directions into a higher-dimensional representation using sinusoidal bases, enabling the network to model high-frequency details. However, under noisy conditions where most high-frequency variations are noise, noticeable “checkerboard effects” are observed in the rendering results (Fig. 7 (a)). This issue arises because P.E. universally assumes that high-frequency details can occur at any position , an assumption valid for clean RGB supervision but leading to overfitting on noise when optimized on noisy RGB.

Refer to caption
Figure 8. Clean RGB, NIR, and noisy RGB in frequency domain. represents the energy ratio of high/low frequency.

Fig. 8 visualizes the frequency-domain distributions of clean RGB, NIR, and noisy RGB. The clean RGB and NIR show similar frequency distributions indicating smooth, natural structures, whereas the noisy RGB exhibits strong dispersion across the entire frequency range dominated by random high-frequency noise. This observation is quantified by the high-to-low frequency energy ratio . The noisy RGB exhibits a much larger ratio (), while clean RGB and NIR are significantly lower ( and , respectively), indicating that noisy RGB is dominated by high-frequency noise. This suggests that under the supervision of extremely noisy RGB, traditional positional encoding is prone to overfit meaningless high-frequency noise, resulting in “checkerboard effects.”

Here the insight is to fully utilize the similarity between NIR and RGB in the frequency domain. Fig. 8 suggests that clean RGB possesses a similar frequency distribution to NIR, with slightly richer high-frequency content due to color/texture variations. Based on this analysis, instead of applying P.E. to , the encoding is applied to the NIR estimation , while are directly fed into the RGB MLP. Concretely, let denote the standard sinusoidal positional encoding; the RGB MLP takes , , and as input and predicts the RGB color :

(1)

This strategy acts as a frequency-aligned modulation that amplifies structurally consistent frequencies present in NIR while suppressing noise-driven high-frequency components from the noisy RGB supervision. Fig. 7 (b) shows that this crucial modification effectively eliminates checkerboard artifacts and improves structural reconstruction.

3.4. Resolving the RGB-NIR Ambiguities

Despite the strong structural guidance provided by NIR, several “blending” effects appear in the rendering results (Fig. 6 (a)), hindering accurate recovery of RGB colors and texture details. This issue arises from the inherent NIR-to-RGB color ambiguity, where a single NIR value can correspond to multiple RGB values. As illustrated in the zoomed-in region of Fig. 6, points with identical NIR values may exhibit distinct RGB colors, making the learning process ill-posed for MLPs. Consequently, the network fails to distinguish these variations and instead produces blended RGB estimates for points sharing the same NIR value (Fig. 6 (a)). Although the RGB MLP takes 3D coordinates and 2D view directions as additional inputs, the smoothness of MLPs prevents the model from producing high-frequency RGB estimates for spatially adjacent points with similar NIR values, resulting in soft transition artifacts.

Refer to caption
Figure 9. Distribution of RGB regarding to the NIR value.
Refer to caption
Figure 10. Qualitative comparison results on synthetic dataset.
Refer to caption
Figure 11. Qualitative ablation results on synthetic dataset. (a) Result w/o NIR-P.E. (b) Result w/o C.C. MLP. (c) Result with full components. Zoom in for the best view.
媒体内容 · 前往原文查看
Table 1. Quantitative comparison results on synthetic dataset. Best and second best results are annotated with bold and underline.
Methods
SSIM PSNR LPIPS SSIM PSNR LPIPS SSIM PSNR LPIPS SSIM PSNR LPIPS SSIM PSNR LPIPS
Restormer 0.6577 17.60 0.3561 0.5886 16.29 0.5397 0.5489 15.43 0.6779 0.4993 14.91 0.8265 0.4941 14.43 0.8245
ScaleMap 0.6902 17.49 0.2793 0.5874 15.98 0.4611 0.4973 15.10 0.6552 0.3749 14.50 0.8393 0.2873 13.91 0.9317
NVEU 0.4586 15.44 0.5025 0.4373 15.07 0.5938 0.4055 14.06 0.6709 0.3552 12.38 0.7253 0.3108 11.26 0.7596
SANet 0.7556 17.52 0.2966 0.6431 15.53 0.4376 0.4934 14.62 0.6785 0.2366 14.21 1.1068 0.1592 13.64 1.2321
NAID 0.5351 16.13 0.3939 0.5223 14.85 0.6353 0.5061 14.17 0.7686 0.4973 13.80 0.8144 0.4929 13.61 0.8257
RawNeRF 0.7047 17.49 0.2384 0.6244 15.99 0.3678 0.5730 15.15 0.5391 0.5371 14.67 0.7377 0.5157 14.21 0.8651
LLNeRF 0.5844 17.82 0.6202 0.5023 13.96 0.7477 0.4569 11.87 0.8321 0.4240 10.12 0.8910 0.4045 9.01 0.9193
Ours 0.7572 19.53 0.2361 0.7441 19.52 0.2354 0.7485 19.70 0.2365 0.7299 19.87 0.2485 0.7028 19.68 0.2670

To address this issue, a Color Code MLP (C.C. MLP) is introduced to learn a color distribution for each NIR value . Specifically, the C.C. MLP predicts a non-uniform log probability distribution , which is used to obtain the one-hot color code :

(2)

where Gumbel-Softmax (Jang et al., 2016) is introduced to maintain differentiability. are i.i.d. samples from the distribution (Gumbel, 1954; Maddison et al., 2014; Jang et al., 2016), and denotes the predicted log probabilities for each of the categories. The final output of the C.C. MLP is a -dimensional one-hot vector , which is then used as input to the RGB MLP for color prediction. is set to by default. This design enables the model to resolve NIR-to-RGB ambiguity by learning a probabilistic mapping from NIR values to plausible RGB variations. As illustrated in Fig. 6, the incorporation of the C.C. MLP significantly improves rendering quality by addressing the fundamental NIR-to-RGB ambiguity.

To further demonstrate the effectiveness of the proposed C.C. MLP, the RGB feature distributions of the GT, the prediction with C.C. MLP, and the prediction without C.C. MLP are visualized in Fig. 9. The RGB values are projected onto a 1D space using PCA, and points are randomly sampled for comparison. The GT distribution exhibits three separate clusters corresponding to distinct RGB colors (circled in ①, ②, and ③). While the prediction without C.C. MLP correctly estimates the RGB distribution within ①, it fails to distinguish clusters ② and ③, as they share similar NIR values. In contrast, the model equipped with C.C. MLP successfully separates these clusters, indicating its ability to resolve the NIR-to-RGB ambiguity and preserve diverse color representations. Fig. 5 shows the final architecture.

3.5. Optimization

Loss function. All MLPs are jointly optimized by minimizing the L2 photometric loss between the rendered NIR estimation and RGB estimation and their corresponding NIR observation and noisy RGB observation :

(3)

where is a hyperparameter to amplify the exposure of .

Training details. The model is optimized for iterations using the Adam optimizer (Kingma and Ba, 2014) with a learning rate of , , and . The training process requires approximately 3 hours on one NVIDIA RTX 4090 GPU, with a peak memory consumption of 9 GB.

4. Evaluations

4.1. Benchmarks

Refer to caption
Figure 12. Qualitative ablation results on real-world captures. (a) Result w/o NIR-P.E. (b) Result w/o C.C. MLP. (c) Result with full components. Zoom in for the best view.
Refer to caption
Figure 13. Visual comparisons on real-world captures. Zoom in for the best view.

Synthetic data. Evaluating the proposed method requires a multi-view dataset containing low-light RGB and corresponding NIR images. Since no such dataset is publicly available, a synthetic dataset is generated using Mitsuba3 (Jakob et al., 2022). Each scene consists of randomly sampled camera views, rendering both a normal RGB image and a NIR image at a resolution of . To simulate noisy RGB images, low-light conditions are created by scaling the RGB pixel values by . Physics-based shot noise and read noise are then synthesized following (Brooks et al., 2019). The dataset contains scenes, each with scale factors representing different noise levels. For each scene, views are reserved for testing, and the remaining views are used for training.

媒体内容 · 前往原文查看
Table 2. Quantitative ablation results on synthetic dataset.
Methods Metrics
w/o NIR-P.E. SSIM 0.7563 0.7365 0.7476 0.7292 0.6936
PSNR 19.48 19.49 19.62 19.75 19.69
LPIPS 0.2378 0.2360 0.2438 0.2657 0.3072
w/o C.C. MLP SSIM 0.7428 0.7294 0.7324 0.7153 0.6923
PSNR 19.34 19.25 19.42 19.61 19.42
LPIPS 0.2563 0.2573 0.2566 0.2614 0.2775
Ours SSIM 0.7572 0.7441 0.7485 0.7299 0.7028
PSNR 19.53 19.52 19.70 19.87 19.68
LPIPS 0.2361 0.2354 0.2365 0.2485 0.2670

Real-world data. To further evaluate the applicability of the proposed method in real-world scenarios, real-world scenes are captured using a JAI FS-3200T10GE-NNC camera under low-light conditions. Each scene contains views with a resolution of . An NIR LED is used for illumination. For each scene, views are reserved for testing, and the remaining views are used for training. Comparisons are first performed in sRGB space against various SOTA baselines. Raw data is converted into 8-bit sRGB space using RawPy. Evaluations are then conducted in 16-bit RAW space, demonstrating the generality of the proposed method. For results on sRGB space, manual color correction are applied on the output for visualization. All results are reported based on the test views. COLMAP (Schönberger and Frahm, 2016) is used to calibrate the camera poses.

4.2. Synthetic Data Evaluation

Comparison. Although this study focuses on extremely noisy scenarios under low-light conditions, synthetic experiments are conducted across multiple noise levels, ranging from moderate to severe noise interference in extreme darkness. The proposed model is compared with SOTAs from different domains, covering both supervised and unsupervised models: Restormer (Zamir et al., 2022) trains a transformer for RGB image denoising. ScaleMap (Yan et al., 2013) performs RGB-NIR denoising with hand-crafted operators. NVEU (Niu et al., 2023c) leverages large-scale unpaired clean RGB to train an RGB-NIR dark denoising model in an unsupervised way. SANet (Sheng et al., 2023) and NAID (Xu et al., 2024) are supervised models for NIR–RGB dark denoising. RawNeRF (Mildenhall et al., 2022) enhances multi-view RAW images. LLNeRF (Wang et al., 2023) is designed for multi-view sRGB low-light enhancement. The results are reported in Tab. 1 and Fig. 10. The proposed model consistently produces accurate results without using clean RGB as supervision, even in the most challenging scenarios.

In addition, a two-stage baseline is implemented where NIR and RGB images are first fused using a 2D fusion method such as SANet (Sheng et al., 2023), NAID (Xu et al., 2024), and NVEU (Niu et al., 2023c), followed by training a NeRF on the fused results. The quantitative results are reported in Tab. 3. The proposed model achieves superior performance across different noise levels. Qualitative comparisons are provided in the supplementary, showing that the two-stage pipeline struggles to restore accurate RGB, whereas the proposed method maintains high fidelity. This performance gap arises because 2D fusion models cannot exploit multi-view consistency when combining RGB and NIR modalities. The proposed approach achieves robustness across different noise levels without using clean RGB for optimization.

Ablation Study. In addition to the analysis in Sec. 3, key components of the proposed model are further ablated to evaluate their effectiveness. w/o NIR-P.E.: A variant that removes NIR-based positional encoding. w/o C.C. MLP: A variant that excludes the Color Code MLP (C.C. MLP) from the architecture. The corresponding qualitative and quantitative results are presented in Fig. 11 and Tab. 2, respectively. The results indicate that removing NIR-based positional encoding leads to pronounced “checkerboard effects,” resulting in visually unsatisfactory renderings. Similarly, removing the C.C. MLP causes the model to misestimate RGB values for spatially adjacent 3D points sharing identical NIR values, leading to noticeable blending artifacts.

4.3. Real-World Data Evaluation

媒体内容 · 前往原文查看
Table 3. Quantitative Results of “2D Fusion + NeRF” baseline under different scale factor .
Methods Metrics
NVEU + NeRF SSIM 0.5278 0.5076 0.4644 0.4183 0.3610
PSNR 16.03 15.71 14.58 12.89 11.90
LPIPS 0.4652 0.5571 0.6339 0.6946 0.7214
SANet + NeRF SSIM 0.7539 0.6553 0.5834 0.5468 0.5231
PSNR 17.54 15.46 14.66 14.63 14.42
LPIPS 0.2878 0.4299 0.5979 0.7612 0.8674
NAID + NeRF SSIM 0.6102 0.5658 0.4506 0.3871 0.3619
PSNR 16.71 16.14 14.42 13.01 12.23
LPIPS 0.4478 0.6391 0.7793 0.8398 0.8623
Ours SSIM 0.7572 0.7441 0.7485 0.7299 0.7028
PSNR 19.53 19.52 19.70 19.87 19.68
LPIPS 0.2361 0.2354 0.2365 0.2485 0.2670
Refer to caption
Figure 14. Qualitative Comparison results in RAW space. Zoom in for the best view.

To evaluate the applicability of different methods, experiments are conducted on real-world scenes. For quantitative assessment, commonly used non-reference image quality metrics are employed, including Perceptual Index (PI) (Gu et al., 2022), MUSIQ (Ke et al., 2021), and MANIQA (Yang et al., 2022). In addition, a Human Subjective Evaluation (HSE) is conducted to directly reflect human perceptual judgment. Specifically, ten volunteers are asked to independently evaluate the results and assign an integer score from 1 (poor) to 5 (excellent).

Ablation Study. The qualitative ablation results on real-world captures are presented in Fig. 12. The model exhibits “checkerboard effects” when NIR-P.E. is removed. Additionally, excluding the C.C. MLP leads to blending artifacts in local regions where pixels share identical NIR values but differ in RGB colors (e.g., the eyes of the owl), resulting in visually degraded RGB restorations. Corresponding quantitative results are reported in Tab. 4. The model with all components achieves the best performance across all metrics and obtains the best scores in human evaluations.

媒体内容 · 前往原文查看
Table 4. Real-world quantitative ablation results. Best and second best results are annotated with bold and underline.
Models PI MUSIQ MANIQA HSE
w/o NIR-P.E. 5.531 49.17 0.2734 3.150
w/o C.C. MLP 5.586 55.73 0.2406 3.175
Ours 5.415 56.51 0.2752 3.450
媒体内容 · 前往原文查看
Table 5. Real-world comparisons with “2D Fusion + NeRF”.
Models PI MUSIQ MANIQA HSE
NVEU + NeRF 7.746 33.74 0.2571 3.050
SANet + NeRF 9.571 28.27 0.2655 3.075
NAID + NeRF 9.359 24.96 0.2694 2.950
Ours 5.415 56.51 0.2752 3.450

Comparison. The comparison results on real-world captures are shown in Fig. 13 and Tab. 6. Existing methods exhibit limited generalization capability in real-world settings, often leading to texture loss and color distortions. In general, the proposed method achieves superior results, preserving both structural integrity and color fidelity across all scenes. Additional qualitative comparisons, including those with the “2D Fusion + NeRF” baseline, are provided in the supplementary. The proposed model achieves more accurate RGB recovery, effectively restoring both appearance and geometric details. Quantitative comparison results are presented in Tab. 5, where the proposed method achieves the best performance. The proposed method is further evaluated in the RAW domain to demonstrate its generality. Specifically, the 12-bit RAW Bayer data is linearly normalized by , without applying any white-balance or color-correction operations. Comparisons are conducted with RawNeRF (Mildenhall et al., 2022), NVEU (Niu et al., 2023c), and NAID (Xu et al., 2024). For NVEU and NAID, the two green channels are averaged, and the resulting RGB values are stacked into a three-channel image. Qualitative and quantitative results are reported in Fig. 14 and Tab. 7. The results indicate that: (1) the proposed method generalizes effectively to the RAW domain, and (2) it generally achieves superior performance compared to RawNeRF, NVEU, and NAID. White balance postprocessing is not applied for RAW data visualization in Fig. 14 to provide a direct and unbiased visualization.

4.4. Discussions and Limitations

The model is limited to static scenes, which constrains its applicability in dynamic environments. Also, when the RGB and NIR are completely unaligned (e.g., using one RGB camera and another NIR camera for free capture), accurate camera poses for RGB observation are hard to obtain. We leave these for future work.

媒体内容 · 前往原文查看
Table 6. Real-world quantitative comparison results.
Models PI MUSIQ MANIQA HSE
Restormer 8.639 16.08 0.2333 2.550
ScaleMap 5.776 29.61 0.1706 2.875
NVEU 4.817 36.01 0.2639 3.025
SANet 5.843 35.03 0.1993 3.075
NAID 9.408 19.98 0.2732 2.925
LLNeRF 9.721 13.23 0.2618 2.625
Ours 5.415 56.51 0.2752 3.450
媒体内容 · 前往原文查看
Table 7. Quantitative comparison results in RAW space. Best and second best results are annotated with bold and underline.
Models PI MUSIQ MANIQA HSE
NVEU 4.821 46.71 0.2482 2.850
NAID 9.924 34.76 0.2240 2.625
RawNeRF 6.783 46.78 0.2251 3.150
Ours 5.874 55.70 0.2527 3.775

5. Conclusion

This paper introduces a new 3D-aware model for RGB–NIR dark imaging. Built upon the volume-rendering framework, a novel implicit neural fusion architecture is carefully designed with several effective components. Both synthetic and real-world datasets are provided to demonstrate the superiority and robustness of the proposed model across various scenarios. Without supervision from clean RGB data, the method achieves performance exceeding state-of-the-art approaches, generalizing across different noise levels without using the clean RGB for supervision.

Acknowledgments

This work was supported in part by JSPS KAKENHI Grant Number 24KK0209, the Forest Digital Twin Project under the Partnership Agreement for Social Value Creation between UTokyo and SMBC Group, the Hokkaido Sarabetsu Village ”Endowed Chair for Field Phenomics” projects in Japan, and the Advanced AI Talent Development to Lead the Next-Generation AI for Intelligent Society (BOOST NAIS) of The University of Tokyo.

References

  • S. Anwar and N. Barnes (2019) Real image denoising with feature attention. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 3155–3164. Cited by: §1, §2.
  • T. Brooks, B. Mildenhall, T. Xue, J. Chen, D. Sharlet, and J. T. Barron (2019) Unprocessing images for learned raw denoising. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11036–11045. Cited by: §4.1.
  • A. Buades, B. Coll, and J. Morel (2005) A non-local algorithm for image denoising. In 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05), Vol. 2, pp. 60–65. Cited by: §1.
  • L. Chen, X. Chu, X. Zhang, and J. Sun (2022) Simple baselines for image restoration. In European conference on computer vision, pp. 17–33. Cited by: §1, §2.
  • K. Dabov, A. Foi, V. Katkovnik, and K. Egiazarian (2007) Image denoising by sparse 3-d transform-domain collaborative filtering. IEEE Transactions on image processing 16 (8), pp. 2080–2095. Cited by: §1.
  • X. Deng and P. L. Dragotti (2020) Deep convolutional neural network for multi-modal image restoration and fusion. IEEE transactions on pattern analysis and machine intelligence 43 (10), pp. 3333–3348. Cited by: §2.
  • W. Dong, G. Shi, and X. Li (2012a) Nonlocal image restoration with bilateral variance estimation: a low-rank approach. IEEE transactions on image processing 22 (2), pp. 700–711. Cited by: §1, §2.
  • W. Dong, L. Zhang, G. Shi, and X. Li (2012b) Nonlocally centralized sparse representation for image restoration. IEEE transactions on Image Processing 22 (4), pp. 1620–1630. Cited by: §1, §2.
  • E. Eisemann and F. Durand (2004) Flash photography enhancement via intrinsic relighting. ACM transactions on graphics (TOG) 23 (3), pp. 673–678. Cited by: §2.
  • J. Gu, H. Cai, C. Dong, J. S. Ren, R. Timofte, Y. Gong, S. Lao, S. Shi, J. Wang, S. Yang, et al. (2022) NTIRE 2022 challenge on perceptual image quality assessment. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 951–967. Cited by: §4.3.
  • S. Gu, L. Zhang, W. Zuo, and X. Feng (2014) Weighted nuclear norm minimization with application to image denoising. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2862–2869. Cited by: §1.
  • E. J. Gumbel (1954) Statistical theory of extreme values and some practical applications: a series of lectures. Vol. 33, US Government Printing Office. Cited by: §3.4.
  • K. He, J. Sun, and X. Tang (2012) Guided image filtering. IEEE transactions on pattern analysis and machine intelligence 35 (6), pp. 1397–1409. Cited by: §2.
  • W. Jakob, S. Speierer, N. Roussel, M. Nimier-David, D. Vicini, T. Zeltner, B. Nicolet, M. Crespo, V. Leroy, and Z. Zhang (2022) Mitsuba 3 renderer Note: https://mitsuba-renderer.org Cited by: §4.1.
  • E. Jang, S. Gu, and B. Poole (2016) Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144. Cited by: §3.4.
  • X. Jia, X. Wei, X. Cao, and H. Foroosh (2019) Comdefend: an efficient image compression model to defend adversarial examples. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 6084–6092. Cited by: §1.
  • S. Jin, B. Yu, M. Jing, Y. Zhou, J. Liang, and R. Ji (2022) Darkvisionnet: low-light imaging via rgb-nir fusion with deep inconsistency prior. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36, pp. 1104–1112. Cited by: §1, §2.
  • J. Ke, Q. Wang, Y. Wang, P. Milanfar, and F. Yang (2021) Musiq: multi-scale image quality transformer. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 5148–5157. Cited by: §4.3.
  • D. P. Kingma and J. Ba (2014) Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980. Cited by: §3.5.
  • D. Krishnan and R. Fergus (2009) Dark flash photography. ACM Trans. Graph. 28 (3), pp. 96. Cited by: §2.
  • Y. Li, J. Huang, N. Ahuja, and M. Yang (2019) Joint image filtering with deep convolutional networks. IEEE transactions on pattern analysis and machine intelligence 41 (8), pp. 1909–1923. Cited by: §2.
  • C. J. Maddison, D. Tarlow, and T. Minka (2014) A* sampling. Advances in neural information processing systems 27. Cited by: §3.4.
  • B. Mildenhall, P. Hedman, R. Martin-Brualla, P. P. Srinivasan, and J. T. Barron (2022) Nerf in the dark: high dynamic range view synthesis from noisy raw images. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 16190–16199. Cited by: §2, §3.1, §4.2, §4.3.
  • B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng (2021) Nerf: representing scenes as neural radiance fields for view synthesis. Communications of the ACM 65 (1), pp. 99–106. Cited by: §2.
  • M. Niu, T. Chen, Y. Zhan, Z. Li, X. Ji, and Y. Zheng (2024) Rs-nerf: neural radiance fields from rolling shutter images. In European Conference on Computer Vision, pp. 163–180. Cited by: §2.
  • M. Niu, Z. Li, Y. Zhan, H. H. Nguyen, I. Echizen, and Y. Zheng (2023a) Physics-based adversarial attack on near-infrared human detector for nighttime surveillance camera systems. In Proceedings of the 31st ACM International Conference on Multimedia, pp. 8799–8807. Cited by: §2.
  • M. Niu, Z. Li, Z. Zhong, and Y. Zheng (2023b) Visibility constrained wide-band illumination spectrum design for seeing-in-the-dark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13976–13985. Cited by: §1.
  • M. Niu, Y. Zhan, Q. Zhu, Z. Li, W. Wang, Z. Zhong, X. Sun, and Y. Zheng (2026) Motion-aware animatable gaussian avatars deblurring. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 40140–40151. Cited by: §1.
  • M. Niu, Z. Zhong, and Y. Zheng (2023c) NIR-assisted video enhancement via unpaired 24-hour data. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10778–10788. Cited by: §4.2, §4.2, §4.3.
  • G. Oh, J. Back, J. Heo, and B. Moon (2023) Robust image denoising of no-flash images guided by consistent flash images. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, pp. 1993–2001. Cited by: §2.
  • N. Pearl, T. Treibitz, and S. Korman (2022) Nan: noise-aware nerfs for burst-denoising. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12672–12681. Cited by: §2, §3.1.
  • G. Petschnigg, R. Szeliski, M. Agrawala, M. Cohen, H. Hoppe, and K. Toyama (2004) Digital photography with flash and no-flash image pairs. ACM transactions on graphics (TOG) 23 (3), pp. 664–672. Cited by: §2.
  • J. L. Schönberger and J. Frahm (2016) Structure-from-motion revisited. In Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §4.1.
  • Z. Sheng, Z. Yu, X. Liu, S. Cao, Y. Liu, H. Shen, and H. Zhang (2023) Structure aggregation for cross-spectral stereo image guided denoising. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13997–14006. Cited by: §1, §2, §4.2, §4.2.
  • R. Wan, B. Shi, W. Yang, B. Wen, L. Duan, and A. C. Kot (2022) Purifying low-light images via near-infrared enlightened image. IEEE Transactions on Multimedia 25, pp. 8006–8019. Cited by: §2.
  • H. Wang, X. Xu, K. Xu, and R. W. Lau (2023) Lighting up nerf via unsupervised decomposition and enhancement. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 12632–12641. Cited by: §2, §3.1, §4.2.
  • Q. Wang, Y. Cui, Y. Li, Y. Ruan, B. Zhu, and W. Ren (2024) RFFNet: towards robust and flexible fusion for low-light image denoising. In Proceedings of the 32nd ACM International Conference on Multimedia, pp. 836–845. Cited by: §1, §2.
  • Y. Wang, H. Wang, L. Wang, X. Wang, L. Zhu, W. Lu, and H. Huang (2025) Complementary advantages: exploiting cross-field frequency correlation for nir-assisted image denoising. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 12679–12689. Cited by: §2.
  • Z. Wang, X. Cun, J. Bao, W. Zhou, J. Liu, and H. Li (2022) Uformer: a general u-shaped transformer for image restoration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 17683–17693. Cited by: §1, §2.
  • Z. Xia, M. Gharbi, F. Perazzi, K. Sunkavalli, and A. Chakrabarti (2021) Deep denoising of flash and no-flash pairs for photography in low-light environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2063–2072. Cited by: §2.
  • J. Xiong, J. Wang, W. Heidrich, and S. Nayar (2021) Seeing in extra darkness using a deep-red flash. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10000–10009. Cited by: §1, §2.
  • J. Xu, L. Zhang, and D. Zhang (2018) A trilateral weighted sparse coding scheme for real-world image denoising. In Proceedings of the European conference on computer vision (ECCV), pp. 20–36. Cited by: §1, §2.
  • R. Xu, Z. Zhang, R. Wu, and W. Zuo (2024) NIR-assisted image denoising: a selective fusion approach and a real-world benchmark dataset. IEEE Transactions on Multimedia. Cited by: §1, §2, §4.2, §4.2, §4.3.
  • Q. Yan, X. Shen, L. Xu, S. Zhuo, X. Zhang, L. Shen, and J. Jia (2013) Cross-field joint image restoration via scale map. In Proceedings of the IEEE International Conference on Computer Vision, pp. 1537–1544. Cited by: §2, §4.2.
  • S. Yang, T. Wu, S. Shi, S. Lao, Y. Gong, M. Cao, J. Wang, and Y. Yang (2022) Maniqa: multi-dimension attention network for no-reference image quality assessment. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 1191–1200. Cited by: §4.3.
  • S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, M. Yang, and L. Shao (2021) Multi-stage progressive image restoration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 14821–14831. Cited by: §1, §2.
  • S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M. Yang (2022) Restormer: efficient transformer for high-resolution image restoration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 5728–5739. Cited by: §1, §2, §4.2.
  • Y. Zhan, Z. Li, M. Niu, Z. Zhong, S. Nobuhara, K. Nishino, and Y. Zheng (2024) Kfd-nerf: rethinking dynamic nerf with kalman filter. In European Conference on Computer Vision, pp. 1–18. Cited by: §2.
  • R. Zhang, Z. Yu, Z. Sheng, J. Ying, S. Cao, S. Chen, B. Yang, J. Li, and H. Shen (2025) SGDFormer: one-stage transformer-based architecture for cross-spectral stereo image guided denoising. Information Fusion 113, pp. 102603. Cited by: §2.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org