Muyao Niu
The University of Tokyo
Tokyo
Japan
muyao.niu@gmail.com
Mingze Ma
Adelaide University
Adelaide
Australia
mingze.ma@adelaide.edu.au
Yifan Zhan
The University of Tokyo
Tokyo
Japan
zhan-yifan@g.ecc.u-tokyo.ac.jp
Qingtian Zhu
The University of Tokyo
Tokyo
Japan
qtzhu@g.ecc.u-tokyo.ac.jp
Zhihang Zhong
School of Artificial Intelligence (SAI)
Shanghai
zhongzhihang@sjtu.edu.cn
Wei Guo
The University of Tokyo
Tokyo
Japan
guowei@g.ecc.u-tokyo.ac.jp
Chang Wen Chen
Hong Kong Polytechnic University
Hong Kong SAR
changwen.chen@polyu.edu.hk
Yinqiang Zheng
The University of Tokyo
Tokyo
Japan
yqzheng@ai.u-tokyo.ac.jp
Abstract.
Robust low-light imaging remains challenging for the community. Recent studies have explored fusing Near-Infrared (NIR) with noisy RGB to achieve improved enhancement, yet most methods depend on carefully curated training data pairs, with limited robustness under different scenarios. This paper offers a new perspective for RGB-NIR low-light imaging by incorporating 3D-aware neural modeling. Without using clean RGB supervision, a powerful model can be optimized to implicitly fuse extremely noisy RGB observations with NIR cues in 3D space, effectively recovering clean RGB images. The proposed model obviates the requirement for clean RGB data collection, generalizes across different noise levels. Extensive evaluations on synthetic and real data demonstrate its superiority. Codes available: https://github.com/MyNiuuu/3DarkFusion
1. Introduction
Robust dark imaging remains a fundamental challenge in computer vision. Despite rapid advances in camera hardware, noise interference under extreme low-light conditions continues to degrade image quality. Conventional strategies, such as long-exposure and burst photography, are prone to motion blur (Niu et al., 2026) and visibility issues (Xiong et al., 2021; Niu et al., 2023b), limiting their practicality.
To mitigate these issues, numerous low-light enhancement methods have been proposed, ranging from traditional denoising operators (Buades et al., 2005; Dabov et al., 2007; Dong et al., 2012b; Jia et al., 2019; Xu et al., 2018; Dong et al., 2012a; Gu et al., 2014) to modern deep learning–based networks (Zamir et al., 2021; Anwar and Barnes, 2019; Chen et al., 2022; Wang et al., 2022; Zamir et al., 2022). Among these, several image- and video-based RGB–NIR low-light enhancement methods have recently attracted increasing attention (Jin et al., 2022; Sheng et al., 2023; Wang et al., 2024; Xu et al., 2024) due to the invisibility of near-infrared signals and their complementary structural cues in darkness. However, existing approaches still face critical limitations. Supervised models are typically trained on “noisy RGB–NIR–clean RGB” triplets from specific domains, resulting in poor cross-domain generalization and limited robustness under varying noise levels. This leads to a natural question: Can robust low-light enhancement be achieved using only NIR and noisy RGB observations, without relying on any form of clean RGB supervision?
This paper investigates a novel perspective on this question by leveraging a multi-view, 3D-aware neural implicit model for efficient RGB–NIR dark imaging without relying on clean RGB supervision. Building upon the volume rendering framework and the analysis of several intuitive baseline architectures, a 3D-aware neural implicit fusion architecture is carefully redesigned to fully exploit the potential of RGB-NIR modalities. In addition, an NIR-modulated positional encoding mechanism is introduced to address the inherent limitations of conventional positional encoding under severe noise, effectively suppressing noise-induced overfitting. Finally, a Color Code MLP is presented to resolve the fundamentally ill-posed mapping between NIR and RGB through a learned color code distribution. Comprehensive evaluations are conducted on both synthetic and real-world multi-view datasets containing NIR images and noisy RGB observations. The proposed framework demonstrates strong performance under extreme low-light conditions compared with alternative approaches. Moreover, it generalizes across varying noise levels compared to existing methods, without using the clean RGB for supervision. Codes and data will be released to support future research. The key contributions of this work are: 1) A new 3D-aware fusion model is proposed for RGB-NIR dark imaging. Different from existing RGB-NIR models, it effectively fuses noisy RGB observation with NIR cues in 3D space without requiring clean RGB supervision. 2) Based on the characteristic of RGB-NIR modalities, two insightful components are introduced to unlock the complementary potentials, further improving performance. 3) Both synthetic and real-world datasets are contributed, demonstrating the superiority of the proposed model across various scenarios.
2. Related Work
Low-light photography. Early low-light imaging methods rely on handcrafted techniques (Dong et al., 2012b; Xu et al., 2018; Dong et al., 2012a), while recent methods (Zamir et al., 2021; Anwar and Barnes, 2019; Chen et al., 2022; Wang et al., 2022) achieve superior performance with deep learning. For example, DnCNN (Zamir et al., 2021) employs CNN for denoising, while transformer-based methods like Restormer (Zamir et al., 2022) further improve the quality.
Denoising with multi-view models. Since the introduction of Neural Radiance Fields (NeRFs) (Mildenhall et al., 2021; Niu et al., 2024; Zhan et al., 2024), several studies (Mildenhall et al., 2022; Pearl et al., 2022; Wang et al., 2023) have investigated their potential for image denoising. RawNeRF (Mildenhall et al., 2022) performs novel-view synthesis directly on RAW data. LLNeRF (Wang et al., 2023) explores denoising in the sRGB domain. Despite these advancements, leveraging NIR for robust multi-view low-light imaging remains unexplored.
Guided image denoising. Beyond approaches that rely solely on RGB information for denoising, several studies (Eisemann and Durand, 2004; Krishnan and Fergus, 2009; Petschnigg et al., 2004; He et al., 2012; Yan et al., 2013; Deng and Dragotti, 2020; Li et al., 2019; Oh et al., 2023; Xia et al., 2021; Xiong et al., 2021; Jin et al., 2022; Sheng et al., 2023; Wang et al., 2024; Zhang et al., 2025; Xu et al., 2024) have explored the use of additional modalities, such as burst flash (Eisemann and Durand, 2004; Krishnan and Fergus, 2009; Petschnigg et al., 2004; He et al., 2012; Oh et al., 2023; Xiong et al., 2021) and Near-Infrared (NIR) imaging (Yan et al., 2013; Jin et al., 2022; Sheng et al., 2023; Wang et al., 2024; Xu et al., 2024; Wan et al., 2022; Wang et al., 2025; Niu et al., 2023a). SANet (Sheng et al., 2023) estimates a clean structure map for RGB-NIR fusion. NAID (Xu et al., 2024) further proposes a Selective Fusion Module.
3. Approach
3.1. Preliminary: MLPs and Volume Rendering
Previous research (Mildenhall et al., 2022; Pearl et al., 2022; Wang et al., 2023) indicates that integrating Multi-Layer Perceptrons (MLPs) with volume rendering can aid in image denoising. This ability arises from two factors: 1) MLPs inherently favor smooth, low-frequency predictions, and 2) the volume-rendering process enforces 3D consistency by aggregating information across multiple viewpoints. To validate this behavior, a vanilla NeRF architecture from (Mildenhall et al., 2022) was trained using noisy RGB inputs. The results (Fig. 4) show that although the method reduces noise to some extent, it introduces significant artifacts. While the volume-rendering framework remains robust and crucial with the enforcement of multi-view consistency, the characteristics of MLPs need to be carefully handled to fully exploit the potentials of RGB–NIR signals, particularly when supervised by noisy RGB data. The following sections retain the standard volume-rendering formulation without modification, while introducing a completely redesigned MLP architecture specifically tailored to the complementary properties of RGB–NIR modalities.
3.2. Fusing NIR and RGB in 3D Space
Intuitive parallel structure. A straightforward parallel architecture jointly optimizes two separate MLPs for RGB and NIR observations. Since the primary purpose of incorporating NIR images is to provide structural guidance, the NIR MLP is used to estimate the volume density , while a gradient-stopping strategy is applied to prevent degraded RGB observations from adversely affecting its learning process. The architectural design and corresponding qualitative results are presented in Fig. 2 (a) and Fig. 3 (a), respectively. The additional structural information from NIR substantially improves the reconstruction of overall scene geometry compared with the vanilla NeRF MLP (Fig. 4). However, the model still fails to recover detailed 2D texture information already present in the NIR images, such as the background texture and the floor in the “Lego” scene. This limitation arises because the RGB MLP does not directly exploit the predicted NIR values, which encode rich texture information crucial for high-fidelity reconstruction.
NIR-conditioned structure. Building on the previous observations, the architecture is improved by feeding the NIR output into the RGB MLP while ensuring that gradients from the RGB MLP do not propagate back to the NIR MLP. This modification enables the RGB MLP to leverage the structural guidance provided by NIR while maintaining the integrity of the NIR representation. The revised architecture and its corresponding qualitative results are shown in Fig. 2 (b) and Fig. 3 (b), respectively. This enhanced design substantially improves both 3D structural reconstruction and the recovery of fine 2D texture details. By implicitly fusing NIR predictions with RGB predictions in 3D space, the model effectively exploits additional structural information for improved results.
3.3. Revisiting Positional Encoding with NIR
Positional Encoding (P.E.) transforms 3D coordinates and 2D view directions into a higher-dimensional representation using sinusoidal bases, enabling the network to model high-frequency details. However, under noisy conditions where most high-frequency variations are noise, noticeable “checkerboard effects” are observed in the rendering results (Fig. 7 (a)). This issue arises because P.E. universally assumes that high-frequency details can occur at any position , an assumption valid for clean RGB supervision but leading to overfitting on noise when optimized on noisy RGB.
Fig. 8 visualizes the frequency-domain distributions of clean RGB, NIR, and noisy RGB. The clean RGB and NIR show similar frequency distributions indicating smooth, natural structures, whereas the noisy RGB exhibits strong dispersion across the entire frequency range dominated by random high-frequency noise. This observation is quantified by the high-to-low frequency energy ratio . The noisy RGB exhibits a much larger ratio (), while clean RGB and NIR are significantly lower ( and , respectively), indicating that noisy RGB is dominated by high-frequency noise. This suggests that under the supervision of extremely noisy RGB, traditional positional encoding is prone to overfit meaningless high-frequency noise, resulting in “checkerboard effects.”
Here the insight is to fully utilize the similarity between NIR and RGB in the frequency domain. Fig. 8 suggests that clean RGB possesses a similar frequency distribution to NIR, with slightly richer high-frequency content due to color/texture variations. Based on this analysis, instead of applying P.E. to , the encoding is applied to the NIR estimation , while are directly fed into the RGB MLP. Concretely, let denote the standard sinusoidal positional encoding; the RGB MLP takes , , and as input and predicts the RGB color :
| (1) |
This strategy acts as a frequency-aligned modulation that amplifies structurally consistent frequencies present in NIR while suppressing noise-driven high-frequency components from the noisy RGB supervision. Fig. 7 (b) shows that this crucial modification effectively eliminates checkerboard artifacts and improves structural reconstruction.
3.4. Resolving the RGB-NIR Ambiguities
Despite the strong structural guidance provided by NIR, several “blending” effects appear in the rendering results (Fig. 6 (a)), hindering accurate recovery of RGB colors and texture details. This issue arises from the inherent NIR-to-RGB color ambiguity, where a single NIR value can correspond to multiple RGB values. As illustrated in the zoomed-in region of Fig. 6, points with identical NIR values may exhibit distinct RGB colors, making the learning process ill-posed for MLPs. Consequently, the network fails to distinguish these variations and instead produces blended RGB estimates for points sharing the same NIR value (Fig. 6 (a)). Although the RGB MLP takes 3D coordinates and 2D view directions as additional inputs, the smoothness of MLPs prevents the model from producing high-frequency RGB estimates for spatially adjacent points with similar NIR values, resulting in soft transition artifacts.
| Methods | |||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| SSIM | PSNR | LPIPS | SSIM | PSNR | LPIPS | SSIM | PSNR | LPIPS | SSIM | PSNR | LPIPS | SSIM | PSNR | LPIPS | |
| Restormer | 0.6577 | 17.60 | 0.3561 | 0.5886 | 16.29 | 0.5397 | 0.5489 | 15.43 | 0.6779 | 0.4993 | 14.91 | 0.8265 | 0.4941 | 14.43 | 0.8245 |
| ScaleMap | 0.6902 | 17.49 | 0.2793 | 0.5874 | 15.98 | 0.4611 | 0.4973 | 15.10 | 0.6552 | 0.3749 | 14.50 | 0.8393 | 0.2873 | 13.91 | 0.9317 |
| NVEU | 0.4586 | 15.44 | 0.5025 | 0.4373 | 15.07 | 0.5938 | 0.4055 | 14.06 | 0.6709 | 0.3552 | 12.38 | 0.7253 | 0.3108 | 11.26 | 0.7596 |
| SANet | 0.7556 | 17.52 | 0.2966 | 0.6431 | 15.53 | 0.4376 | 0.4934 | 14.62 | 0.6785 | 0.2366 | 14.21 | 1.1068 | 0.1592 | 13.64 | 1.2321 |
| NAID | 0.5351 | 16.13 | 0.3939 | 0.5223 | 14.85 | 0.6353 | 0.5061 | 14.17 | 0.7686 | 0.4973 | 13.80 | 0.8144 | 0.4929 | 13.61 | 0.8257 |
| RawNeRF | 0.7047 | 17.49 | 0.2384 | 0.6244 | 15.99 | 0.3678 | 0.5730 | 15.15 | 0.5391 | 0.5371 | 14.67 | 0.7377 | 0.5157 | 14.21 | 0.8651 |
| LLNeRF | 0.5844 | 17.82 | 0.6202 | 0.5023 | 13.96 | 0.7477 | 0.4569 | 11.87 | 0.8321 | 0.4240 | 10.12 | 0.8910 | 0.4045 | 9.01 | 0.9193 |
| Ours | 0.7572 | 19.53 | 0.2361 | 0.7441 | 19.52 | 0.2354 | 0.7485 | 19.70 | 0.2365 | 0.7299 | 19.87 | 0.2485 | 0.7028 | 19.68 | 0.2670 |
To address this issue, a Color Code MLP (C.C. MLP) is introduced to learn a color distribution for each NIR value . Specifically, the C.C. MLP predicts a non-uniform log probability distribution , which is used to obtain the one-hot color code :
| (2) |
where Gumbel-Softmax (Jang et al., 2016) is introduced to maintain differentiability. are i.i.d. samples from the distribution (Gumbel, 1954; Maddison et al., 2014; Jang et al., 2016), and denotes the predicted log probabilities for each of the categories. The final output of the C.C. MLP is a -dimensional one-hot vector , which is then used as input to the RGB MLP for color prediction. is set to by default. This design enables the model to resolve NIR-to-RGB ambiguity by learning a probabilistic mapping from NIR values to plausible RGB variations. As illustrated in Fig. 6, the incorporation of the C.C. MLP significantly improves rendering quality by addressing the fundamental NIR-to-RGB ambiguity.
To further demonstrate the effectiveness of the proposed C.C. MLP, the RGB feature distributions of the GT, the prediction with C.C. MLP, and the prediction without C.C. MLP are visualized in Fig. 9. The RGB values are projected onto a 1D space using PCA, and points are randomly sampled for comparison. The GT distribution exhibits three separate clusters corresponding to distinct RGB colors (circled in ①, ②, and ③). While the prediction without C.C. MLP correctly estimates the RGB distribution within ①, it fails to distinguish clusters ② and ③, as they share similar NIR values. In contrast, the model equipped with C.C. MLP successfully separates these clusters, indicating its ability to resolve the NIR-to-RGB ambiguity and preserve diverse color representations. Fig. 5 shows the final architecture.
3.5. Optimization
Loss function. All MLPs are jointly optimized by minimizing the L2 photometric loss between the rendered NIR estimation and RGB estimation and their corresponding NIR observation and noisy RGB observation :
| (3) |
where is a hyperparameter to amplify the exposure of .
Training details. The model is optimized for iterations using the Adam optimizer (Kingma and Ba, 2014) with a learning rate of , , and . The training process requires approximately 3 hours on one NVIDIA RTX 4090 GPU, with a peak memory consumption of 9 GB.
4. Evaluations
4.1. Benchmarks
Synthetic data. Evaluating the proposed method requires a multi-view dataset containing low-light RGB and corresponding NIR images. Since no such dataset is publicly available, a synthetic dataset is generated using Mitsuba3 (Jakob et al., 2022). Each scene consists of randomly sampled camera views, rendering both a normal RGB image and a NIR image at a resolution of . To simulate noisy RGB images, low-light conditions are created by scaling the RGB pixel values by . Physics-based shot noise and read noise are then synthesized following (Brooks et al., 2019). The dataset contains scenes, each with scale factors representing different noise levels. For each scene, views are reserved for testing, and the remaining views are used for training.
| Methods | Metrics | |||||
|---|---|---|---|---|---|---|
| w/o NIR-P.E. | SSIM | 0.7563 | 0.7365 | 0.7476 | 0.7292 | 0.6936 |
| PSNR | 19.48 | 19.49 | 19.62 | 19.75 | 19.69 | |
| LPIPS | 0.2378 | 0.2360 | 0.2438 | 0.2657 | 0.3072 | |
| w/o C.C. MLP | SSIM | 0.7428 | 0.7294 | 0.7324 | 0.7153 | 0.6923 |
| PSNR | 19.34 | 19.25 | 19.42 | 19.61 | 19.42 | |
| LPIPS | 0.2563 | 0.2573 | 0.2566 | 0.2614 | 0.2775 | |
| Ours | SSIM | 0.7572 | 0.7441 | 0.7485 | 0.7299 | 0.7028 |
| PSNR | 19.53 | 19.52 | 19.70 | 19.87 | 19.68 | |
| LPIPS | 0.2361 | 0.2354 | 0.2365 | 0.2485 | 0.2670 |
Real-world data. To further evaluate the applicability of the proposed method in real-world scenarios, real-world scenes are captured using a JAI FS-3200T10GE-NNC camera under low-light conditions. Each scene contains views with a resolution of . An NIR LED is used for illumination. For each scene, views are reserved for testing, and the remaining views are used for training. Comparisons are first performed in sRGB space against various SOTA baselines. Raw data is converted into 8-bit sRGB space using RawPy. Evaluations are then conducted in 16-bit RAW space, demonstrating the generality of the proposed method. For results on sRGB space, manual color correction are applied on the output for visualization. All results are reported based on the test views. COLMAP (Schönberger and Frahm, 2016) is used to calibrate the camera poses.
4.2. Synthetic Data Evaluation
Comparison. Although this study focuses on extremely noisy scenarios under low-light conditions, synthetic experiments are conducted across multiple noise levels, ranging from moderate to severe noise interference in extreme darkness. The proposed model is compared with SOTAs from different domains, covering both supervised and unsupervised models: Restormer (Zamir et al., 2022) trains a transformer for RGB image denoising. ScaleMap (Yan et al., 2013) performs RGB-NIR denoising with hand-crafted operators. NVEU (Niu et al., 2023c) leverages large-scale unpaired clean RGB to train an RGB-NIR dark denoising model in an unsupervised way. SANet (Sheng et al., 2023) and NAID (Xu et al., 2024) are supervised models for NIR–RGB dark denoising. RawNeRF (Mildenhall et al., 2022) enhances multi-view RAW images. LLNeRF (Wang et al., 2023) is designed for multi-view sRGB low-light enhancement. The results are reported in Tab. 1 and Fig. 10. The proposed model consistently produces accurate results without using clean RGB as supervision, even in the most challenging scenarios.
In addition, a two-stage baseline is implemented where NIR and RGB images are first fused using a 2D fusion method such as SANet (Sheng et al., 2023), NAID (Xu et al., 2024), and NVEU (Niu et al., 2023c), followed by training a NeRF on the fused results. The quantitative results are reported in Tab. 3. The proposed model achieves superior performance across different noise levels. Qualitative comparisons are provided in the supplementary, showing that the two-stage pipeline struggles to restore accurate RGB, whereas the proposed method maintains high fidelity. This performance gap arises because 2D fusion models cannot exploit multi-view consistency when combining RGB and NIR modalities. The proposed approach achieves robustness across different noise levels without using clean RGB for optimization.
Ablation Study. In addition to the analysis in Sec. 3, key components of the proposed model are further ablated to evaluate their effectiveness. w/o NIR-P.E.: A variant that removes NIR-based positional encoding. w/o C.C. MLP: A variant that excludes the Color Code MLP (C.C. MLP) from the architecture. The corresponding qualitative and quantitative results are presented in Fig. 11 and Tab. 2, respectively. The results indicate that removing NIR-based positional encoding leads to pronounced “checkerboard effects,” resulting in visually unsatisfactory renderings. Similarly, removing the C.C. MLP causes the model to misestimate RGB values for spatially adjacent 3D points sharing identical NIR values, leading to noticeable blending artifacts.
4.3. Real-World Data Evaluation
| Methods | Metrics | |||||
|---|---|---|---|---|---|---|
| NVEU + NeRF | SSIM | 0.5278 | 0.5076 | 0.4644 | 0.4183 | 0.3610 |
| PSNR | 16.03 | 15.71 | 14.58 | 12.89 | 11.90 | |
| LPIPS | 0.4652 | 0.5571 | 0.6339 | 0.6946 | 0.7214 | |
| SANet + NeRF | SSIM | 0.7539 | 0.6553 | 0.5834 | 0.5468 | 0.5231 |
| PSNR | 17.54 | 15.46 | 14.66 | 14.63 | 14.42 | |
| LPIPS | 0.2878 | 0.4299 | 0.5979 | 0.7612 | 0.8674 | |
| NAID + NeRF | SSIM | 0.6102 | 0.5658 | 0.4506 | 0.3871 | 0.3619 |
| PSNR | 16.71 | 16.14 | 14.42 | 13.01 | 12.23 | |
| LPIPS | 0.4478 | 0.6391 | 0.7793 | 0.8398 | 0.8623 | |
| Ours | SSIM | 0.7572 | 0.7441 | 0.7485 | 0.7299 | 0.7028 |
| PSNR | 19.53 | 19.52 | 19.70 | 19.87 | 19.68 | |
| LPIPS | 0.2361 | 0.2354 | 0.2365 | 0.2485 | 0.2670 |
To evaluate the applicability of different methods, experiments are conducted on real-world scenes. For quantitative assessment, commonly used non-reference image quality metrics are employed, including Perceptual Index (PI) (Gu et al., 2022), MUSIQ (Ke et al., 2021), and MANIQA (Yang et al., 2022). In addition, a Human Subjective Evaluation (HSE) is conducted to directly reflect human perceptual judgment. Specifically, ten volunteers are asked to independently evaluate the results and assign an integer score from 1 (poor) to 5 (excellent).
Ablation Study. The qualitative ablation results on real-world captures are presented in Fig. 12. The model exhibits “checkerboard effects” when NIR-P.E. is removed. Additionally, excluding the C.C. MLP leads to blending artifacts in local regions where pixels share identical NIR values but differ in RGB colors (e.g., the eyes of the owl), resulting in visually degraded RGB restorations. Corresponding quantitative results are reported in Tab. 4. The model with all components achieves the best performance across all metrics and obtains the best scores in human evaluations.
| Models | PI | MUSIQ | MANIQA | HSE |
|---|---|---|---|---|
| w/o NIR-P.E. | 5.531 | 49.17 | 0.2734 | 3.150 |
| w/o C.C. MLP | 5.586 | 55.73 | 0.2406 | 3.175 |
| Ours | 5.415 | 56.51 | 0.2752 | 3.450 |
| Models | PI | MUSIQ | MANIQA | HSE |
|---|---|---|---|---|
| NVEU + NeRF | 7.746 | 33.74 | 0.2571 | 3.050 |
| SANet + NeRF | 9.571 | 28.27 | 0.2655 | 3.075 |
| NAID + NeRF | 9.359 | 24.96 | 0.2694 | 2.950 |
| Ours | 5.415 | 56.51 | 0.2752 | 3.450 |
Comparison. The comparison results on real-world captures are shown in Fig. 13 and Tab. 6. Existing methods exhibit limited generalization capability in real-world settings, often leading to texture loss and color distortions. In general, the proposed method achieves superior results, preserving both structural integrity and color fidelity across all scenes. Additional qualitative comparisons, including those with the “2D Fusion + NeRF” baseline, are provided in the supplementary. The proposed model achieves more accurate RGB recovery, effectively restoring both appearance and geometric details. Quantitative comparison results are presented in Tab. 5, where the proposed method achieves the best performance. The proposed method is further evaluated in the RAW domain to demonstrate its generality. Specifically, the 12-bit RAW Bayer data is linearly normalized by , without applying any white-balance or color-correction operations. Comparisons are conducted with RawNeRF (Mildenhall et al., 2022), NVEU (Niu et al., 2023c), and NAID (Xu et al., 2024). For NVEU and NAID, the two green channels are averaged, and the resulting RGB values are stacked into a three-channel image. Qualitative and quantitative results are reported in Fig. 14 and Tab. 7. The results indicate that: (1) the proposed method generalizes effectively to the RAW domain, and (2) it generally achieves superior performance compared to RawNeRF, NVEU, and NAID. White balance postprocessing is not applied for RAW data visualization in Fig. 14 to provide a direct and unbiased visualization.
4.4. Discussions and Limitations
The model is limited to static scenes, which constrains its applicability in dynamic environments. Also, when the RGB and NIR are completely unaligned (e.g., using one RGB camera and another NIR camera for free capture), accurate camera poses for RGB observation are hard to obtain. We leave these for future work.
| Models | PI | MUSIQ | MANIQA | HSE |
|---|---|---|---|---|
| Restormer | 8.639 | 16.08 | 0.2333 | 2.550 |
| ScaleMap | 5.776 | 29.61 | 0.1706 | 2.875 |
| NVEU | 4.817 | 36.01 | 0.2639 | 3.025 |
| SANet | 5.843 | 35.03 | 0.1993 | 3.075 |
| NAID | 9.408 | 19.98 | 0.2732 | 2.925 |
| LLNeRF | 9.721 | 13.23 | 0.2618 | 2.625 |
| Ours | 5.415 | 56.51 | 0.2752 | 3.450 |
| Models | PI | MUSIQ | MANIQA | HSE |
|---|---|---|---|---|
| NVEU | 4.821 | 46.71 | 0.2482 | 2.850 |
| NAID | 9.924 | 34.76 | 0.2240 | 2.625 |
| RawNeRF | 6.783 | 46.78 | 0.2251 | 3.150 |
| Ours | 5.874 | 55.70 | 0.2527 | 3.775 |
5. Conclusion
This paper introduces a new 3D-aware model for RGB–NIR dark imaging. Built upon the volume-rendering framework, a novel implicit neural fusion architecture is carefully designed with several effective components. Both synthetic and real-world datasets are provided to demonstrate the superiority and robustness of the proposed model across various scenarios. Without supervision from clean RGB data, the method achieves performance exceeding state-of-the-art approaches, generalizing across different noise levels without using the clean RGB for supervision.
Acknowledgments
This work was supported in part by JSPS KAKENHI Grant Number 24KK0209, the Forest Digital Twin Project under the Partnership Agreement for Social Value Creation between UTokyo and SMBC Group, the Hokkaido Sarabetsu Village ”Endowed Chair for Field Phenomics” projects in Japan, and the Advanced AI Talent Development to Lead the Next-Generation AI for Intelligent Society (BOOST NAIS) of The University of Tokyo.
References
- S. Anwar and N. Barnes (2019) Real image denoising with feature attention. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 3155–3164. Cited by: §1, §2.
- T. Brooks, B. Mildenhall, T. Xue, J. Chen, D. Sharlet, and J. T. Barron (2019) Unprocessing images for learned raw denoising. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11036–11045. Cited by: §4.1.
- A. Buades, B. Coll, and J. Morel (2005) A non-local algorithm for image denoising. In 2005 IEEE computer society conference on computer vision and pattern recognition (CVPR’05), Vol. 2, pp. 60–65. Cited by: §1.
- L. Chen, X. Chu, X. Zhang, and J. Sun (2022) Simple baselines for image restoration. In European conference on computer vision, pp. 17–33. Cited by: §1, §2.
- K. Dabov, A. Foi, V. Katkovnik, and K. Egiazarian (2007) Image denoising by sparse 3-d transform-domain collaborative filtering. IEEE Transactions on image processing 16 (8), pp. 2080–2095. Cited by: §1.
- X. Deng and P. L. Dragotti (2020) Deep convolutional neural network for multi-modal image restoration and fusion. IEEE transactions on pattern analysis and machine intelligence 43 (10), pp. 3333–3348. Cited by: §2.
- W. Dong, G. Shi, and X. Li (2012a) Nonlocal image restoration with bilateral variance estimation: a low-rank approach. IEEE transactions on image processing 22 (2), pp. 700–711. Cited by: §1, §2.
- W. Dong, L. Zhang, G. Shi, and X. Li (2012b) Nonlocally centralized sparse representation for image restoration. IEEE transactions on Image Processing 22 (4), pp. 1620–1630. Cited by: §1, §2.
- E. Eisemann and F. Durand (2004) Flash photography enhancement via intrinsic relighting. ACM transactions on graphics (TOG) 23 (3), pp. 673–678. Cited by: §2.
- J. Gu, H. Cai, C. Dong, J. S. Ren, R. Timofte, Y. Gong, S. Lao, S. Shi, J. Wang, S. Yang, et al. (2022) NTIRE 2022 challenge on perceptual image quality assessment. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 951–967. Cited by: §4.3.
- S. Gu, L. Zhang, W. Zuo, and X. Feng (2014) Weighted nuclear norm minimization with application to image denoising. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2862–2869. Cited by: §1.
- E. J. Gumbel (1954) Statistical theory of extreme values and some practical applications: a series of lectures. Vol. 33, US Government Printing Office. Cited by: §3.4.
- K. He, J. Sun, and X. Tang (2012) Guided image filtering. IEEE transactions on pattern analysis and machine intelligence 35 (6), pp. 1397–1409. Cited by: §2.
- W. Jakob, S. Speierer, N. Roussel, M. Nimier-David, D. Vicini, T. Zeltner, B. Nicolet, M. Crespo, V. Leroy, and Z. Zhang (2022) Mitsuba 3 renderer Note: https://mitsuba-renderer.org Cited by: §4.1.
- E. Jang, S. Gu, and B. Poole (2016) Categorical reparameterization with gumbel-softmax. arXiv preprint arXiv:1611.01144. Cited by: §3.4.
- X. Jia, X. Wei, X. Cao, and H. Foroosh (2019) Comdefend: an efficient image compression model to defend adversarial examples. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 6084–6092. Cited by: §1.
- S. Jin, B. Yu, M. Jing, Y. Zhou, J. Liang, and R. Ji (2022) Darkvisionnet: low-light imaging via rgb-nir fusion with deep inconsistency prior. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 36, pp. 1104–1112. Cited by: §1, §2.
- J. Ke, Q. Wang, Y. Wang, P. Milanfar, and F. Yang (2021) Musiq: multi-scale image quality transformer. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 5148–5157. Cited by: §4.3.
- D. P. Kingma and J. Ba (2014) Adam: a method for stochastic optimization. arXiv preprint arXiv:1412.6980. Cited by: §3.5.
- D. Krishnan and R. Fergus (2009) Dark flash photography. ACM Trans. Graph. 28 (3), pp. 96. Cited by: §2.
- Y. Li, J. Huang, N. Ahuja, and M. Yang (2019) Joint image filtering with deep convolutional networks. IEEE transactions on pattern analysis and machine intelligence 41 (8), pp. 1909–1923. Cited by: §2.
- C. J. Maddison, D. Tarlow, and T. Minka (2014) A* sampling. Advances in neural information processing systems 27. Cited by: §3.4.
- B. Mildenhall, P. Hedman, R. Martin-Brualla, P. P. Srinivasan, and J. T. Barron (2022) Nerf in the dark: high dynamic range view synthesis from noisy raw images. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 16190–16199. Cited by: §2, §3.1, §4.2, §4.3.
- B. Mildenhall, P. P. Srinivasan, M. Tancik, J. T. Barron, R. Ramamoorthi, and R. Ng (2021) Nerf: representing scenes as neural radiance fields for view synthesis. Communications of the ACM 65 (1), pp. 99–106. Cited by: §2.
- M. Niu, T. Chen, Y. Zhan, Z. Li, X. Ji, and Y. Zheng (2024) Rs-nerf: neural radiance fields from rolling shutter images. In European Conference on Computer Vision, pp. 163–180. Cited by: §2.
- M. Niu, Z. Li, Y. Zhan, H. H. Nguyen, I. Echizen, and Y. Zheng (2023a) Physics-based adversarial attack on near-infrared human detector for nighttime surveillance camera systems. In Proceedings of the 31st ACM International Conference on Multimedia, pp. 8799–8807. Cited by: §2.
- M. Niu, Z. Li, Z. Zhong, and Y. Zheng (2023b) Visibility constrained wide-band illumination spectrum design for seeing-in-the-dark. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13976–13985. Cited by: §1.
- M. Niu, Y. Zhan, Q. Zhu, Z. Li, W. Wang, Z. Zhong, X. Sun, and Y. Zheng (2026) Motion-aware animatable gaussian avatars deblurring. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 40140–40151. Cited by: §1.
- M. Niu, Z. Zhong, and Y. Zheng (2023c) NIR-assisted video enhancement via unpaired 24-hour data. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10778–10788. Cited by: §4.2, §4.2, §4.3.
- G. Oh, J. Back, J. Heo, and B. Moon (2023) Robust image denoising of no-flash images guided by consistent flash images. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 37, pp. 1993–2001. Cited by: §2.
- N. Pearl, T. Treibitz, and S. Korman (2022) Nan: noise-aware nerfs for burst-denoising. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12672–12681. Cited by: §2, §3.1.
- G. Petschnigg, R. Szeliski, M. Agrawala, M. Cohen, H. Hoppe, and K. Toyama (2004) Digital photography with flash and no-flash image pairs. ACM transactions on graphics (TOG) 23 (3), pp. 664–672. Cited by: §2.
- J. L. Schönberger and J. Frahm (2016) Structure-from-motion revisited. In Conference on Computer Vision and Pattern Recognition (CVPR), Cited by: §4.1.
- Z. Sheng, Z. Yu, X. Liu, S. Cao, Y. Liu, H. Shen, and H. Zhang (2023) Structure aggregation for cross-spectral stereo image guided denoising. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13997–14006. Cited by: §1, §2, §4.2, §4.2.
- R. Wan, B. Shi, W. Yang, B. Wen, L. Duan, and A. C. Kot (2022) Purifying low-light images via near-infrared enlightened image. IEEE Transactions on Multimedia 25, pp. 8006–8019. Cited by: §2.
- H. Wang, X. Xu, K. Xu, and R. W. Lau (2023) Lighting up nerf via unsupervised decomposition and enhancement. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 12632–12641. Cited by: §2, §3.1, §4.2.
- Q. Wang, Y. Cui, Y. Li, Y. Ruan, B. Zhu, and W. Ren (2024) RFFNet: towards robust and flexible fusion for low-light image denoising. In Proceedings of the 32nd ACM International Conference on Multimedia, pp. 836–845. Cited by: §1, §2.
- Y. Wang, H. Wang, L. Wang, X. Wang, L. Zhu, W. Lu, and H. Huang (2025) Complementary advantages: exploiting cross-field frequency correlation for nir-assisted image denoising. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 12679–12689. Cited by: §2.
- Z. Wang, X. Cun, J. Bao, W. Zhou, J. Liu, and H. Li (2022) Uformer: a general u-shaped transformer for image restoration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 17683–17693. Cited by: §1, §2.
- Z. Xia, M. Gharbi, F. Perazzi, K. Sunkavalli, and A. Chakrabarti (2021) Deep denoising of flash and no-flash pairs for photography in low-light environments. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2063–2072. Cited by: §2.
- J. Xiong, J. Wang, W. Heidrich, and S. Nayar (2021) Seeing in extra darkness using a deep-red flash. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 10000–10009. Cited by: §1, §2.
- J. Xu, L. Zhang, and D. Zhang (2018) A trilateral weighted sparse coding scheme for real-world image denoising. In Proceedings of the European conference on computer vision (ECCV), pp. 20–36. Cited by: §1, §2.
- R. Xu, Z. Zhang, R. Wu, and W. Zuo (2024) NIR-assisted image denoising: a selective fusion approach and a real-world benchmark dataset. IEEE Transactions on Multimedia. Cited by: §1, §2, §4.2, §4.2, §4.3.
- Q. Yan, X. Shen, L. Xu, S. Zhuo, X. Zhang, L. Shen, and J. Jia (2013) Cross-field joint image restoration via scale map. In Proceedings of the IEEE International Conference on Computer Vision, pp. 1537–1544. Cited by: §2, §4.2.
- S. Yang, T. Wu, S. Shi, S. Lao, Y. Gong, M. Cao, J. Wang, and Y. Yang (2022) Maniqa: multi-dimension attention network for no-reference image quality assessment. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 1191–1200. Cited by: §4.3.
- S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, M. Yang, and L. Shao (2021) Multi-stage progressive image restoration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 14821–14831. Cited by: §1, §2.
- S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M. Yang (2022) Restormer: efficient transformer for high-resolution image restoration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 5728–5739. Cited by: §1, §2, §4.2.
- Y. Zhan, Z. Li, M. Niu, Z. Zhong, S. Nobuhara, K. Nishino, and Y. Zheng (2024) Kfd-nerf: rethinking dynamic nerf with kalman filter. In European Conference on Computer Vision, pp. 1–18. Cited by: §2.
- R. Zhang, Z. Yu, Z. Sheng, J. Ying, S. Cao, S. Chen, B. Yang, J. Li, and H. Shen (2025) SGDFormer: one-stage transformer-based architecture for cross-spectral stereo image guided denoising. Information Fusion 113, pp. 102603. Cited by: §2.