PixRestore:基于像素扩散 Transformer 的统一图像修复

HuggingFace Daily Papers(社区热门论文)·2026-08-17 08:00·7天前
AI 导读

PixRestore 提出一种免 VAE 的像素扩散 Transformer(DiT),用于统一图像修复,无需 T2I 预训练。该方法利用 DINO 分层特征与自适应层路由提供密集条件引导,并通过 DINO 对抗目标将多步模型微调为单步生成器。仅约 50M 参数、单步推理约 44 ms,PixRestore 在扩散类方法中取得最佳修复质量与最高效率。

HuggingFace Daily Papers(社区热门论文)
55AI 编辑部评分,满分 100

PixRestore:基于像素扩散 Transformer 的统一图像修复

2026-08-17 08:00· 7天前
AI 导读

PixRestore 提出一种免 VAE 的像素扩散 Transformer(DiT),用于统一图像修复,无需 T2I 预训练。该方法利用 DINO 分层特征与自适应层路由提供密集条件引导,并通过 DINO 对抗目标将多步模型微调为单步生成器。仅约 50M 参数、单步推理约 44 ms,PixRestore 在扩散类方法中取得最佳修复质量与最高效率。

Lingchen Sun1,2   Rongyuan Wu1,2   Xiangtao Kong 1,2   Jixin Zhao 2   Qiaosi Yi 1,2   
Yujing Sun 1,2   Shuaizheng Liu 1,2   Zhengqiang Zhang 1,2   Lei Zhang1,2

1 The Hong Kong Polytechnic University 2 OPPO Research Institute

 Equal contribution.   Corresponding author (cslzhang@comp.polyu.edu.hk).

Refer to caption
Figure 1: PixRestore achieves the best overall performance in terms of restoration quality, model size, and inference speed. Top-left: radar charts comparing PSNR, LPIPS, and degradation removal performance across eight degradation types. Top-right: GFLOPs versus inference latency, where the bubble size denotes the number of parameters. Bottom: visual comparisons on eight restoration tasks. With only about 50M parameters and single-step inference, PixRestore achieves the best overall restoration quality while being the most efficient among diffusion-based methods.

KEYWORDS : Image Restoration, Pixel Diffusion, Degradation-Aware, DINO

 1  Introduction

Image restoration (IR) [1, 2, 3] aims to recover a high-quality (HQ) image from its low-quality (LQ) counterpart corrupted by diverse and often co-occurring degradations such as noise, blur, rain, haze, and low light. Rather than training specialist models per degradation, many efforts have been devoted to pursuing unified image restoration (UIR), i.e., using a single model to handle a broad spectrum of image degradations [4]. Built on CNN- and Transformer-based backbones, conventional regression-based methods [5, 6, 7, 8, 9] are efficient and have made encouraging progress. However, their deterministic L1/L2 objectives and limited capacity tend to yield over-smoothed results and unremoved degradations.

Diffusion models have recently been applied to UIR [10, 8] to improve perceptual quality by leveraging their strong generative capacity. Most recent works finetune pretrained text-to-image (T2I) latent diffusion models, e.g., FoundIR-v2 [11] adapts SDXL [12] with MoE routing and an MLLM [13] captioner, and FLUX-IR [14] finetunes FLUX [15] with reinforced ODE trajectories and cost-aware distillation. T2I pretraining brings rich priors and perceptual realism, but at three costs: (1) Lossy latent bottleneck, as the VAE may discard image textures and details that restoration aims to preserve; (2) Objective mismatch, as T2I priors may synthesize visually plausible but inconsistent details with the input; and (3) Redundant computation, caused by the billion-scale backbones, MLLM planners, MoE routing, VAE coding, and iterative sampling. This motivates us to rethink the suitability of T2I models for UIR. Unlike T2I generation, which synthesizes an image from textual cues, UIR starts from LQ images, which contain rich visual cues. Therefore, UIR requires less open-ended generative capacity than T2I models, but it demands robustness to different degradations and faithful reconstruction of pixel-aligned details.

Motivated by the above observations, we present PixRestore, a VAE-free pixel diffusion transformer (DiT) [16, 17] for UIR. The diffusion backbone is trained from scratch without T2I pretraining. By applying flow matching directly to patchified pixels, PixRestore preserves pixel-aligned evidence while keeping the token sequence tractable. Our method removes the autoencoder overhead and substantially improves content fidelity while maintaining strong perceptual quality. Unlike task-specific IR, UIR must handle diverse and compounded degradations, making a static global degradation representation insufficient. We therefore exploit hierarchical features from the self-supervised visual foundation model DINO [18] to provide adaptive guidance for restoration. Our analysis reveals two properties of DINO features: the layers carry complementary cues, from shallow structures to deep semantics, and their reliability varies across degradation types. Accordingly, we train an adaptive layer router to predict per-layer weights from the LQ input, using LQ–HQ DINO feature similarity as supervision during training. The predicted weights fuse features from more reliable layers into dense conditioning, while less reliable layers receive stronger supervision from the corresponding HQ features.

Finally, to speed up PixRestore at inference time, we first train a multi-step PixRestore model on a large-scale corpus of diverse scenes and degradations, then finetune it into a single-step generator with DINO-based adversarial objectives. Fig. 1 compares PixRestore against existing diffusion-based UIR methods in terms of restoration quality, model size, and inference latency. With only about 50M parameters and single-step inference, PixRestore attains the best quality while being the fastest (about 44 ms) and most compact model among the diffusion-based methods. In addition, our experiments show that larger PixRestore variants can further improve restoration quality, confirming the scalability of our design.

  • We propose PixRestore, a pixel-space DiT for UIR. Working on patchified pixels, PixRestore is free of the VAE and T2I priors, improving image fidelity, perceptual quality, and model efficiency.

  • We introduce adaptive hierarchical visual guidance, providing dense conditioning to guide restoration and supervision to stabilize training under various degradations.

  • We finetune the multi-step model into a single-step generator via DINO-based adversarial objectives, achieving efficient inference with little quality loss.

  • Extensive experiments show that PixRestore achieves superior fidelity, perceptual quality, and robustness on public benchmarks and real-world test sets.

 2  Related Work

Regression-based Unified Restoration. Conventional IR methods are typically developed for a specific degradation, such as noise [1], blur [3], rain [19], haze [2], etc. UIR instead seeks a single model for multiple degradations. Existing UIR methods improve degradation adaptivity with learned degradation representations [5], prompts [6], or expert routing [20], etc. Methods such as PromptIR [6], AirNet [5], and their successors show that degradation-aware modulation can substantially improve multi-task compatibility. However, these models are usually optimized as deterministic LQ-to-HQ regressors with L1/L2 losses, which favor conditional averages, often suppressing high-frequency details and limiting perceptual realism. Their task-level prompts or routing decisions may also generalize poorly to more complex real-world degradations. We instead model restoration as a conditional pixel-space flow and derive dense, per-image guidance from hierarchical DINO features.

Generative Unified Restoration. Generative UIR methods synthesize HQ images conditioned on LQ inputs. One line of research learns the conditional generative process from scratch within the restoration task, keeping the model restoration-native [21, 10]. For example, DiffUIR [10] and DA-CLIP [8] design restoration-specific diffusion pipelines or degradation-aware conditioning to improve fidelity across different degradations. Another line adapts pretrained T2I latent diffusion models with degradation predictors [22], multimodal prompts [23], or routing mechanisms [11]. These methods benefit from strong generative priors, but suffer from the conflict between open-ended image synthesis and faithful restoration. In addition, the VAE compresses the input image before diffusion, removing restoration-sensitive details such as small structures, text strokes, and sharp edges. Large T2I backbones and auxiliary planners or expert modules further increase computational cost. In contrast, our proposed PixRestore operates in a patchified pixel space and uses a scalable transformer with layer-adaptive visual conditioning for efficient and faithful UIR.

Pixel Generative Modeling. Diffusion is originally formulated in pixel space, whereas latent diffusion becomes dominant for the reduced generation costs [24]. Recent work has revisited VAE-free generation using improved DiT architectures [17] and loss functions [25]. For UIR, the LQ image provides dense spatial correspondence to the desired output, but compressing it using a VAE can destroy useful information for restoration. Pixel-space modeling preserves that evidence, but introduces computational challenges due to the long spatial sequence. We address this trade-off by using patchification. Different from previous pixel generators developed for text-conditioned synthesis, PixRestore combines pixel-space flow modeling with degradation-aware hierarchical DINO guidance for unified restoration.

Refer to caption
Figure 2: Overview of PixRestore. A frozen vision encoder extracts multi-layer features from the LQ image, and an adaptive layer router predicts per-layer weights to fuse them into a single conditioning feature. The LQ image and the noisy state xt are patchified, then processed by N DiT blocks, and finally decoded into the HQ output.

 3  Method

Let yhq[1,1]3×H×W be an HQ image and ylq its LQ counterpart corrupted by degradations such as noise, blur, haze, rain, or low light. UIR aims to learn a single model that maps ylq to an estimate y^hq, without any task-specific expert. We formulate UIR as a conditional flow matching problem in pixel space and present PixRestore, whose network framework is shown in Fig. 2. A VAE-free pixel DiT learns the conditional flow directly on RGB pixels, avoiding the lossy compression of a latent autoencoder (Sec. 3.1). To capture both degradation and semantic cues, we use a vision encoder to extract multi-layer dense features from the LQ image, and use an adaptive layer router to predict per-layer weights pl. These weights fuse the features into a single representation, which is injected into the DiT blocks by cross-attention. In addition, the predicted weights enable hierarchical visual supervision, detailed in Sec. 3.2. Finally, for efficient inference, we finetune a single-step generator from the multi-step model via DINO-based adversarial objectives [26] (Sec. 3.3).

3.1  Pixel-space Restoration Diffusion Model

Instead of encoding images into a compressed latent space [23], PixRestore operates directly on RGB pixels. Given the HQ target yhq, we adopt the linear interpolation path xt=(1t)yhq+tϵ with ϵ𝒩(0,I) and t𝒰(0,1). The restoration DiT fθ predicts the clean image as:

y^hq=fθ([ylq;xt],t,(ylq)), (1)

where [;] denotes channel-wise concatenation and (ylq) denotes the multi-layer DINO features of the LQ input. Following JiT [17], the flow-matching objective is defined on the velocity. Therefore, we recover the velocity from the output by vt=(xtyhq)/t and v^t=(xty^hq)/t. To prevent division-by-zero as t0, we clip the denominator of 1/t (by default, at 0.05) during computation. The flow matching loss is flow=v^tvt22.

媒体内容 · 前往原文查看
Table 1: Latent diffusion vs. pixel diffusion under the same UIR training/test setting. The results are averaged over 8 restoration tasks. Pixel-space modeling provides a better overall trade-off in fidelity, perceptual quality, parameter count, and inference speed.
Model VAE Params(M) Inf Time-DM (ms/step) Inf Time-VAE (ms) PSNR (dB)  LPIPS  MUSIQ 
Latent DiT-S SD2VAE (f8c4, ps1) 106.41 25 78 22.10 0.2483 50.38
Latent DiT-S FluxVAE (f8c16, ps1) 106.59 25 78 22.63 0.2109 50.86
Latent DiT-S QwenVAE (f8c16, ps1) 67.37 25 41 22.80 0.2181 51.87
Pixel DiT-S None (ps8) 23.41 25 0 26.62 0.1593 54.32

The concatenated input [ylq;xt] has six channels. A patch embedding partitions the full-resolution pixel grid into tokens, shown in Fig. 2. In this way, a relatively large patch size keeps the token sequence tractable without a VAE. Each DiT block consists of RMSNorm, QK-normalized attention, and rotary positional embeddings [27]. A single shared timestep block produces AdaLN modulation parameters for all Transformer blocks [28]. Multi-layer DINO features are injected via cross-attention in each block. During inference, the model integrates the predicted velocity from Gaussian noise using an Euler solver.

The overall training objective combines the flow loss with two auxiliary terms from the DINO module:

=flow+λwpredwpred+λfeatfeat, (2)

where wpred supervises the adaptive layer router and feat enforces hierarchical feature fidelity. The two losses will be discussed and defined in Sec. 3.2.

Pixel Space vs. Latent Space. Latent DiTs use a VAE to reduce computational complexity before diffusion, but this compression can sacrifice small structures and fine textures. Pixel DiT instead operates on RGB pixels and keeps their spatial structure through patchification, better matching IR tasks, where the output must stay faithful to the input. To verify this, we compare the same DiT-S in the pixel space against latent space built on three widely used VAEs, i.e., SD2VAE [24], FluxVAE [15] and QwenVAE [29], under identical training and test settings, including training data, resolution, diffusion architecture, optimizer, and sampler. The pixel DiT model uses a patch size of 8 to match the latent resolution.

As shown in Table 1, pixel DiT beats all latent DiTs across all metrics (26.62 dB PSNR, 0.1593 LPIPS, 54.32 MUSIQ), significantly outperforming the best baseline with QwenVAE. Removing the VAE also reduces cost and latency (41–78 ms). Using only 23.41M parameters without VAE, our pixel DiT design reproduces more faithful image details, indicating that pixel-space modeling can better match the requirements of UIR. Detailed training and test settings, as well as per-degradation analysis, are provided in the Appendix.

3.2  Adaptive Hierarchical Visual Guidance

Visual Foundation Prior. To handle different types of degradations in the input LQ image, we introduce a frozen vision foundation encoder to provide dense visual cues for UIR. Specifically, we adopt DINOv2 [18] for this purpose because its self-distillation pretraining can produce dense, spatially precise tokens that preserve fine structure and texture while being semantically discriminative. We validate this choice with an experimental comparison against other encoders (CLIP [30], MAE [31], SigLIP [32], DINOv2 [18]) under the same pixel DiT setting, where DINOv2 performs the best on almost all metrics. The experiment details are in the Appendix.

Refer to caption
Figure 3: Motivation of adaptive hierarchical visual guidance. Left: Per-layer DINO feature visualizations for low-light enhancement and de-raindrop. Shallow layers preserve local structures and details, while deeper layers encode global semantics. Right: LQ–HQ feature similarity across DINOv2-B layers for eight types of degradations. We see that different layers are sensitive to different degradations.

Similarity-Guided Adaptive Layer Router. As illustrated in Fig. 3 (left), different DINO layers carry complementary visual cues: shallow layers (l1l2) preserve local structures such as edges and textures for detailed reconstruction, while deeper layers (l8l10) encode global semantics for degradation and content discrimination. Layer sensitivity is also degradation-dependent. As shown in Fig. 3 (right), which presents the LQ-HQ feature similarity, no single layer is sensitive to all degradations. In particular, shallow layers are sensitive to rain, blur, and snow degradations; deeper layers are sensitive to noise; and middle layers are sensitive to low-light, SR, and haze. Using a fixed layer or a uniform average over layers is thus suboptimal. We therefore train a lightweight module to predict per-image layer weights, supervised by paired LQ-HQ similarity.

For each layer l, we measure how much the LQ features retain the HQ content by averaging a cosine similarity and a normalized L2-distance similarity over the projected LQ and HQ patch tokens: sl=12(slcos+sldist). A larger sl means that layer l is more reliable under the observed degradation, so it should have a larger weight:

ql=exp(sl)kexp(sk). (3)

Computing ql requires the HQ image, which is unavailable at inference time. Thus, we train a lightweight predictor ρψ to estimate the weights from the LQ image and features: pl=softmax(ρψ([ylq,Ul])),l, where Ul denotes the projected LQ features of different DINO layers. The cross-entropy loss is used to supervise the training:

wpred=lqllogpl. (4)

The predictor thus learns which DINO layers are more reliable for each input, instead of relying on a uniform mixture.

Adaptive Conditioning and Hierarchical Supervision. The projected LQ features are fused using the predicted weights as follows:

Ufuse=lplUl, (5)

which are then injected into each DiT block via cross-attention, with image tokens as queries and Ufuse as keys and values. To emphasize layers where the LQ features differ most from the HQ features, we further introduce a hierarchical feature supervision loss:

feat=lrllfeat,rl=exp((1sl))kexp((1sk)), (6)

where lfeat is the cosine similarity loss between the HQ features and the restored-output features at layer l. The weights ql (see Eq.  (3)) and rl are complementary, i.e., ql selects reliable content features for conditioning, while rl focuses supervision on layers that need stronger restoration. With wpred and feat defined above, the complete multi-step objective is given by Eq. (2), where both auxiliary weights are set to 0.5, and all layer features are channel-wise RMS normalized before projection.

3.3  Single-step Finetuning

The multi-step PixRestore model requires iterative sampling. We therefore finetune it into a single-step generator for efficient UIR. We initialize the student from the pretrained multi-step teacher and fix the flow time to t=1 so that the generator can predict the clean image in one forward pass from pure Gaussian noise ϵ, i.e., y^hq=fθ([ylq;ϵ],t=1,(ylq)).

Since one-step generation may lose fine textures, we add a DINO-based adversarial objective. Reusing the same frozen encoder and layers, a lightweight multi-layer discriminator D distinguishes the restored features (y^hq) from the HQ features (yhq) with an independent head per layer. The discriminator is trained with the standard binary cross-entropy loss to classify Flyhq as real and Fly^hq as fake:

D=1||l[bce(Dl(Flyhq),1)+bce(Dl(Fly^hq),0)]. (7)

The generator is optimized to fool D:

adv=1||lbce(Dl(Fly^hq),1). (8)

The generator and discriminator are updated alternately. With t=1, the single-step objective is:

=flow+λwpredwpred+λfeatfeat+λadvadv, (9)

where all auxiliary weights are set to 0.5. Operating on frozen DINO tokens rather than raw pixels, the discriminator shares the conditioning prior and adds little overhead. During inference, PixRestore restores an image from one noise sample with the LQ condition in a single step.

 4  Experiments

4.1  Experimental Setup

Training and Test Datasets. We build a training corpus of about 2.83M images covering eight restoration tasks (deblur, dehaze, denoise, de-rainstreak, de-raindrop, desnow, low-light enhancement, and super-resolution (SR)), with samples drawn with equal probability during training.

We evaluate PixRestore under two complementary settings. The first uses public benchmarks with paired GT for fidelity and perceptual evaluation: GoPro [33] and UHD-blur [3] (deblur), RESIDE-6K [2] and UHD-Haze [3] (dehaze), DIV2K [34] (Gaussian noise) and PolyU [35] (denoise), RainDS-real [36] and RealRain-1k [37] (de-rainstreak), RainDS-real [36] and UAV-Rain1k [38] (de-raindrop), UHD-LL [39] and LOLdataset [40] (low-light), WeatherBench [41] (desnow), and RealSR [42] and ScreenSR [43] (SR), all center-cropped to 512 for testing. The second is a real-world test set without GT, with 100 LQ images per degradation (deblur, dehaze, de-rainstreak, de-raindrop, desnow, low-light) from diverse sources [44, 45, 46]. Detailed training and testing data are given in the Appendix.

Compared Methods. We compare with representative UIR methods, including the regression-based PromptIR [6] and diffusion-based DA-CLIP [8], DiffUIR [10], FoundIR [9], FoundIR-v2 [11], UniRestore [47], Flux-IR [14] and FAPE-IR [23]. For fair comparison and to isolate the effect of training data, we evaluate both the official checkpoints and retrained versions of the major baselines on our dataset.

PixRestore Model Settings. PixRestore adopts LightningDiT [27] as the DiT backbone and DINOv2 [18] as the vision encoder. Each Transformer block follows the LightningDiT design and contains a multi-head self-attention module, a cross-attention module, and a feed-forward network. A single shared timestep block produces AdaLN modulation parameters for all Transformer blocks [28]. We provide four variants of the model with different backbone sizes. All variants use a patch size of 8 in pixel patchification. During training, the DINOv2 model is frozen and only the diffusion model is trained. Unless otherwise specified, “PixRestore” refers to the PixRestore-S variant in the paper.

  • PixRestore-S uses a LightningDiT-S backbone with a hidden dimension of 384. It contains 12 Transformer blocks, each with 6 attention heads. DINOv2-S is used as the vision encoder.

  • PixRestore-B uses a LightningDiT-B backbone with a hidden dimension of 768. It contains 12 Transformer blocks, each with 12 attention heads. DINOv2-B is used as the vision encoder.

  • PixRestore-L uses a LightningDiT-L backbone with a hidden dimension of 1024. It contains 24 Transformer blocks, each with 16 attention heads. DINOv2-L is used as the vision encoder.

  • PixRestore-XL uses a LightningDiT-XL backbone with a hidden dimension of 1152. It contains 28 Transformer blocks, each with 16 attention heads. Since DINOv2 does not provide an XL version, we use DINOv2-L as the vision encoder.

Training Details. We train the multi-step model for 250K iterations, and then finetune it into a one-step model for an additional 100K iterations, using AdamW with a learning rate of 1×104 and a batch size of 16 on 512×512 image crops across 8 NVIDIA A800 GPUs. DINO features are extracted from six layers evenly distributed throughout the encoder.

Refer to caption
Figure 4: No-reference quality metrics do not reliably reflect degradation removal. Here PixRestore removes the rainstreaks best, yet MUSIQ and AFINE-NR rank it worst, favoring the LQ input and the degradation-preserving output of FoundIR-v2, while our VLM-based DR-Score demonstrates strong alignment with human perceptual judgments.
媒体内容 · 前往原文查看
Table 2: Quantitative comparison on public benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with are retrained using the same training dataset as ours. Metrics: PSNR, SSIM, LPIPS, DISTS, DR-Score (degradation-removal score).
Method De-rainstreak Denoise Deblur De-raindrop
PSNR SSIM LPIPS DISTS DR-Score PSNR SSIM LPIPS DISTS DR-Score PSNR SSIM LPIPS DISTS DR-Score PSNR SSIM LPIPS DISTS DR-Score
PromptIR 24.50 0.7507 0.3529 0.2381 31.25 32.54 0.9029 0.2459 0.1604 55.35 23.82 0.7584 0.3131 0.2195 27.23 18.81 0.6962 0.3390 0.1885 24.86
PromptIR 28.43 0.8389 0.2659 0.1904 49.99 35.45 0.9377 0.1104 0.1065 79.33 29.10 0.8474 0.2051 0.1560 55.56 23.68 0.8005 0.2236 0.1326 51.74
DiffUIR 24.51 0.7618 0.3484 0.2248 40.13 26.61 0.8380 0.3103 0.2056 59.52 27.86 0.8200 0.2258 0.1699 44.09 18.85 0.7004 0.3455 0.1905 26.50
UniRestore 22.50 0.7291 0.4223 0.2689 34.31 31.75 0.8969 0.2336 0.1669 61.20 23.87 0.7205 0.2405 0.1716 46.94 18.50 0.6377 0.4114 0.2285 27.90
DA-CLIP 24.50 0.7560 0.3347 0.2176 43.09 27.35 0.8245 0.2459 0.1764 64.12 27.40 0.8113 0.1641 0.1303 55.00 20.12 0.6988 0.2715 0.1512 55.20
DA-CLIP 31.61 0.8606 0.1048 0.0885 78.72 33.97 0.8728 0.1462 0.1166 75.14 28.29 0.8228 0.1474 0.1148 58.38 23.58 0.7916 0.1325 0.0852 79.64
FoundIR 26.87 0.8263 0.2454 0.1799 52.37 32.46 0.7925 0.2794 0.1641 55.52 27.29 0.8046 0.2430 0.1781 39.12 18.87 0.6972 0.3586 0.2017 26.21
FoundIR 32.36 0.8888 0.1593 0.1203 71.12 36.16 0.9415 0.1051 0.1245 80.86 29.14 0.8466 0.1986 0.1503 49.86 24.40 0.8218 0.1942 0.1146 65.00
Flux-IR 20.98 0.6233 0.4625 0.2861 26.04 25.71 0.6842 0.4197 0.2367 54.86 23.51 0.6745 0.2746 0.1851 58.70 18.94 0.6459 0.3224 0.1774 39.72
Flux-IR 21.01 0.6344 0.4511 0.2822 31.04 26.80 0.7410 0.3789 0.2051 37.05 22.15 0.6446 0.2818 0.2036 73.18 17.80 0.5785 0.3266 0.2075 71.26
FoundIR-v2 23.17 0.6652 0.3659 0.2218 57.26 26.87 0.7481 0.2785 0.1946 74.47 24.40 0.7079 0.2203 0.1566 73.23 19.63 0.5459 0.2792 0.1563 55.48
FoundIR-v2 27.85 0.7598 0.1773 0.1322 81.03 28.17 0.7497 0.2889 0.1890 75.01 24.98 0.7351 0.1914 0.1323 77.12 20.81 0.5461 0.2300 0.1294 80.28
FAPE-IR 27.51 0.8226 0.2319 0.1679 71.09 32.99 0.9088 0.1238 0.1113 79.31 26.80 0.7895 0.2098 0.1556 50.72 21.39 0.6822 0.2219 0.1316 71.77
FAPE-IR 31.91 0.8695 0.0903 0.0762 84.64 34.22 0.9172 0.0741 0.0750 82.56 27.96 0.8207 0.1577 0.1094 66.75 23.88 0.7223 0.1555 0.0967 81.91
PixRestore 32.28 0.8847 0.0902 0.0905 82.75 34.87 0.9336 0.0624 0.0735 81.89 28.32 0.8284 0.1201 0.0940 70.19 24.48 0.7755 0.1258 0.0882 82.20
PixRestore-B 32.85 0.8898 0.0767 0.0817 84.72 34.62 0.9356 0.0564 0.0706 83.70 29.23 0.8511 0.1051 0.0835 72.69 25.21 0.7976 0.1087 0.0763 84.77
Method Desnow Dehaze Low-light enhancement Super-resolution
PSNR SSIM LPIPS DISTS DR-Score PSNR SSIM LPIPS DISTS DR-Score PSNR SSIM LPIPS DISTS DR-Score PSNR SSIM LPIPS DISTS DR-Score
PromptIR 22.26 0.7939 0.2452 0.1672 24.35 21.34 0.8803 0.1426 0.1012 61.68 10.49 0.4781 0.5410 0.3765 27.87 24.26 0.7372 0.4394 0.2510 33.48
PromptIR 29.32 0.8532 0.1818 0.1402 72.39 21.25 0.8931 0.1348 0.0885 67.89 17.91 0.6881 0.3564 0.2512 58.58 27.59 0.7968 0.2839 0.2208 55.70
DiffUIR 22.95 0.7948 0.2392 0.1667 25.00 20.41 0.8656 0.1640 0.1175 57.19 21.72 0.7082 0.3819 0.2292 62.93 26.54 0.7582 0.3866 0.2378 45.81
UniRestore 22.33 0.7863 0.2464 0.1771 26.12 20.11 0.8469 0.2122 0.1359 67.33 10.94 0.5188 0.5006 0.3165 38.38 24.80 0.7543 0.3548 0.2284 60.49
DA-CLIP 23.60 0.7971 0.2221 0.1558 35.53 22.96 0.8751 0.1311 0.0885 70.33 22.25 0.7915 0.2469 0.1622 73.29 23.73 0.6789 0.3772 0.2367 46.43
DA-CLIP 28.31 0.8282 0.1360 0.1021 81.48 21.92 0.8779 0.1459 0.1077 58.58 18.01 0.8185 0.2129 0.1614 72.79 26.49 0.7585 0.2341 0.1763 65.25
FoundIR 23.03 0.7999 0.2406 0.1630 24.68 15.07 0.7901 0.2582 0.1919 33.17 15.34 0.7473 0.3034 0.2132 62.24 25.85 0.7399 0.4280 0.2492 29.70
FoundIR 29.82 0.8678 0.1524 0.1224 73.00 20.99 0.8903 0.1335 0.0962 64.81 23.34 0.9026 0.1792 0.1439 80.41 27.65 0.7952 0.2883 0.2251 64.69
Flux-IR 21.74 0.7231 0.3434 0.2204 26.80 14.73 0.7599 0.2976 0.1986 46.13 18.85 0.7022 0.3557 0.1998 62.16 22.49 0.6541 0.2903 0.2064 77.19
Flux-IR 21.79 0.6831 0.3023 0.1959 51.21 15.96 0.8061 0.2315 0.1578 52.88 18.40 0.7198 0.3386 0.2010 61.45 20.87 0.5853 0.3569 0.2519 74.89
FoundIR-v2 24.72 0.7347 0.2513 0.1671 67.47 19.06 0.7705 0.1928 0.1337 62.79 17.17 0.7448 0.3132 0.2027 72.49 23.85 0.6661 0.2959 0.1974 76.25
FoundIR-v2 26.28 0.7690 0.1824 0.1311 82.52 19.10 0.7645 0.1933 0.1282 63.22 17.04 0.7307 0.3188 0.2101 68.16 23.79 0.6530 0.2493 0.1682 78.87
FAPE-IR 26.02 0.8191 0.1759 0.1189 60.36 25.28 0.9007 0.0960 0.0650 75.76 19.46 0.7629 0.2580 0.1838 58.29 26.55 0.7732 0.2816 0.1977 68.92
FAPE-IR 30.19 0.8676 0.1136 0.0849 83.66 25.34 0.9056 0.0927 0.0645 76.59 26.04 0.9014 0.1370 0.1068 82.00 27.48 0.7843 0.1952 0.1433 77.81
PixRestore 31.26 0.8859 0.0853 0.0669 84.22 25.46 0.9142 0.0896 0.0650 74.86 25.62 0.8894 0.1360 0.1008 81.58 27.01 0.7730 0.1736 0.1358 77.53
PixRestore-B 31.93 0.8959 0.0688 0.0584 84.51 26.64 0.9228 0.0789 0.0574 76.60 25.72 0.8935 0.1298 0.0955 83.08 27.07 0.7773 0.1601 0.1289 77.88

Evaluation. We assess fidelity with PSNR and SSIM [48] and perceptual quality with LPIPS [49] and DISTS [50]. Degradation removal is a key indicator for UIR, reflecting whether a method truly eliminates the target degradation, especially in real-world cases where only no-reference (NR) metrics apply. However, existing NR metrics (e.g., MUSIQ [51] and AFINE-NR [52]) do not measure it well. As shown in Fig. 4, PixRestore removes rainstreaks most effectively but scores the worst on MUSIQ and AFINE-NR, which favor the LQ input and the degradation-preserving output of FoundIR-v2. We therefore propose DR-Score as an auxiliary diagnostic metric. It uses a vision-language model (VLM, e.g., Gemini-3.1 Pro [53]) to judge whether the target degradation has been removed. Because VLM outputs can be stochastic, we test each image five times and report the average score. In the Appendix, we detail the prompt design of DR-Score, and demonstrate its strong alignment with human perceptual judgments.

Refer to caption
Figure 5: Visual comparisons on desnow (top) and low-light enhancement (bottom). PixRestore removes degradations effectively and recovers more faithful details and colors.
媒体内容 · 前往原文查看
Table 3: Complexity comparison of different methods. “NFE” denotes number of evaluations. All values are measured with input 1×3×512×512 on a single NVIDIA A800 GPU, with 5 warmup iterations and averaged over 100 runs.
Method NFE Params (M) FLOPs (G) Latency (ms)
PromptIR - 35.59 1382 334
DiffUIR 4 36.26 3839 777
DA-CLIP 100 231.76 112927 18071
FoundIR 4 36.26 3839 777
UniRestore 1 1071.20 5255 158
FoundIR-v2 20 16910.09 119334 18293
Flux-IR 21 17698.23 795290 5790
FAPE-IR 1 21575.43 64758 1011
PixRestore 1 53.70 658 44
PixRestore-B 1 210.89 1842 79

4.2  Public Benchmark Results

Table 2 compares the competing methods on eight tasks. First, we can see that most of the models retrained on our dataset improve over their original counterparts, showing the effectiveness of our training data, which consist of more samples with diverse content. Second, the two variants of PixRestore show the best overall performance, with its variants ranking first or second on most metrics including reference-based metrics and DR-Score. The results also expose clear differences among methods. The regression model PromptIR remains competitive on denoise, but its performance drops on deblur, rain, haze, low-light, and SR, where the information is lost more severely. Restoration-native diffusion models (DiffUIR, DA-CLIP, FoundIR) generally improve perceptual quality and degradation removal on several tasks. DA-CLIP is relatively strong on de-rainstreak, de-raindrop, and desnow, while FoundIR performs better on dehaze and denoise. This suggests that such models benefit from generative modeling, but their performance varies noticeably across degradations.

Pretrained latent T2I methods show a different trade-off. The SDXL-based FoundIR-v2 attains high DR-Scores but lower fidelity and perceptual performance, which is consistent with the input-detail loss introduced by latent VAE compression. FLUX-based FAPE-IR is among the strongest methods on several degradations (e.g., de-rainstreak and denoise), whereas Flux-IR shows much less consistent performance. This suggests that a stronger generative prior alone does not guarantee better UIR performance. FAPE-IR requires a complex auxiliary design (e.g., Qwen2-VL [54] and SigLIP in FAPE-IR) to adapt the T2I prior to UIR, as reflected by its substantially large parameter count.

PixRestore is more closely aligned with the restoration objective. Flow matching on patchified pixels preserves spatial evidence that a VAE may weaken, while the hierarchical DINO guidance adapts the conditioning to each degradation rather than relying on a single global prior. Fig. 5 shows visual comparisons. On the desnow and low-light enhancement benchmarks, competing methods often leave residual degradations or recover less faithful details, whereas PixRestore removes the degradations more thoroughly and preserves sharper structures. These gains are achieved with substantially lower inference cost, as shown in Table 3.

4.3  Model Complexity Comparisons

Table 3 compares model size, computation, and latency under the same resolution and hardware. With only 53.7M parameters (including the frozen DINO encoder), 658G FLOPs, and 44 ms per image, PixRestore is much lighter than other UIR models, running about 7–23× faster than PromptIR, DiffUIR, FoundIR, and FAPE-IR. Its computation is only about 1/181 of FoundIR-v2 and 1/1209 of Flux-IR. In addition, PixRestore-B further improves UIR performance while still keeping the model compact and efficient, with 210.89M parameters, 1842G FLOPs, and 79 ms latency. Our results suggest that strong UIR models may not require a lossy latent VAE or a massive T2I prior.

媒体内容 · 前往原文查看
Table 4: Ablation studies of PixRestore. “Avg.” denotes uniform averaging of selected DINO features for conditioning or supervision, while “Adap.” denotes adaptive averaging of selected DINO features for conditioning or supervision. “NFE” denotes the number of evaluations.
ID Variant NFE Conditioning Supervision PSNR SSIM LPIPS MUSIQ
A0 Pixel DiT-S 10 None None 26.62 0.8454 0.1593 54.32
A1 + single-layer conditioning 10 layer 2 None 27.07 0.8457 0.1561 54.45
A2 + single-layer conditioning 10 layer 5 None 27.36 0.8487 0.1489 54.64
A3 + single-layer conditioning 10 layer 11 None 27.12 0.8491 0.1508 54.60
A4 + multi-layer conditioning 10 Avg. 2 layers None 27.62 0.8540 0.1412 54.90
A5 + multi-layer conditioning 10 Avg. 6 layers None 27.72 0.8536 0.1407 55.01
A6 + hierarchical loss 10 Avg. 6 layers Avg. 6 layers 27.36 0.8444 0.1239 54.59
A7 + adaptive hierarchical visual guidance 10 Adap. 6 layers Adap. 6 layers 27.66 0.8500 0.1209 54.97
A8 + adaptive hierarchical visual guidance 4 Adap. 6 layers Adap. 6 layers 27.75 0.8584 0.1228 54.07
A9 + adaptive hierarchical visual guidance 1 Adap. 6 layers Adap. 6 layers 28.07 0.8640 0.1202 53.34
A10 + single-step finetuning (PixRestore) 1 Adap. 6 layers Adap. 6 layers 28.49 0.8589 0.1120 55.52
媒体内容 · 前往原文查看
Table 5: Comparison between diffusion pretraining and finetuning with regression training under the same objective and total training iterations.
Scheme Training iterations PSNR SSIM LPIPS MUSIQ
Regression Training 350k 27.00 0.8179 0.1494 52.00
Flow pretraining + one-step finetuning (ours) 250k + 100k 28.49 0.8589 0.1120 55.52

4.4  Ablation Studies

We conduct ablations to verify the main designs of PixRestore. The results are reported in Table 4 and Table 5, which are averaged over 15 public benchmarks covering 8 degradation types. All variants use LightningDiT-S in the pixel space as the baseline.

Effect of DINO Conditioning. Starting from the plain Pixel DiT-S baseline (A0), adding a single DINO feature consistently improves all metrics. The best single-layer choice is layer 5 (A2), which improves PSNR from 26.62 to 27.36, SSIM from 0.8454 to 0.8487, LPIPS from 0.1593 to 0.1489, and MUSIQ from 54.32 to 54.64. This shows that DINO features provide effective guidance for UIR.

Single-layer or Multi-layer Guidance. Using multiple DINO layers is better than using one fixed layer. Compared with the best single-layer setting A2, averaging 2 layers (A4) improves PSNR from 27.36 to 27.62 and LPIPS from 0.1489 to 0.1412. Averaging 6 layers (A5) further raises PSNR to 27.72 and MUSIQ to 55.01. This shows that shallow and deep DINO layers provide complementary cues for UIR.

Effect of Hierarchical Supervision. After introducing the hierarchical feature loss, LPIPS improves clearly from 0.1407 (A5) to 0.1239 (A6), showing better perceptual restoration. However, PSNR drops from 27.72 to 27.36 and SSIM drops from 0.8536 to 0.8444. This suggests that uniform feature-space supervision helps recover more realistic details, but does not give the best overall balance.

Adaptive Hierarchical Visual Guidance. We further compare A6 and A7. Both of them use multi-layer conditioning and supervision, but A6 uses uniform averaging while A7 uses adaptive layer weighting. A7 improves PSNR from 27.36 to 27.66, SSIM from 0.8444 to 0.8500, MUSIQ from 54.59 to 54.97, and LPIPS from 0.1239 to 0.1209. This shows that adaptive guidance better exploits the degradation-dependent reliability of DINO layers.

Different NFE. We further study the effect of reducing the number of function evaluations (NFE). Compared with A7 which uses 10 NFE, A8 with 4 NFE improves PSNR from 27.66 to 27.75 and SSIM from 0.8500 to 0.8584, while LPIPS changes from 0.1209 to 0.1228. When NFE is further reduced to 1, A9 still improves PSNR to 28.07 and SSIM to 0.8640, with LPIPS of 0.1202. MUSIQ drops from 54.97 to 54.07 and 53.34, but the overall results remain competitive. Interestingly, reducing NFE slightly improves PSNR/SSIM in our setting, possibly because fewer Euler updates reduce the accumulation of integration errors and over-smoothing. These results suggest that PixRestore maintains strong restoration quality even in the one-step setting.

Effect of Single-step Finetuning. We finetune the multi-step model into a one-step generator. Compared with A9, the final PixRestore (A10) improves PSNR from 28.07 to 28.49, LPIPS from 0.1202 to 0.1120, and MUSIQ from 53.34 to 55.52, while keeping NFE at 1. This shows that one-step fine-tuning improves restoration quality while retaining one-step efficiency.

Flow Pretraining vs. Regression Training. Table 5 compares our training pipeline with direct regression training under the same backbone, loss functions (with one-step finetuning), and total training iterations. Flow pretraining followed by one-step finetuning improves PSNR from 27.00 to 28.49, SSIM from 0.8179 to 0.8589, LPIPS from 0.1494 to 0.1120, and MUSIQ from 52.00 to 55.52. This shows that the gain comes from the flow-based pretraining stage.

Refer to caption
Figure 6: Scaling behavior of PixRestore under varying backbone GFLOPs and patch sizes.

4.5  Scalability

Fig. 6 illustrates the scalability of PixRestore by varying the Transformer size and pixel patch size. A smaller patch yields more image tokens and thus more computation, which improves LPIPS for the same backbone; for example, the S model improves the LPIPS from about 0.129 with p=16 to about 0.101 with p=4. Enlarging the backbone at a fixed patch size brings a similar gain. Across all configurations, LPIPS decreases as GFLOPs increase, with a correlation of 0.96. Consistent with the scaling behavior reported for DiT [16], this smooth trend suggests that PixRestore can benefit from scaling along two axes: a smaller patch preserves more local evidence, while a larger backbone provides stronger global modeling.

媒体内容 · 前往原文查看
Table 6: Quantitative comparison on real-world test set. The best and second-best results for each metric are highlighted in red bold and blue italic, respectively. Retrained methods are marked with . ‘PR’ denotes the proposed PixRestore, the results of which are shaded in pink.
Degradation Metric PromptIR PromptIR DiffUIR DA-CLIP DA-CLIP FoundIR FoundIR UniRestore FoundIR-v2 FoundIR-v2 Flux-IR Flux-IR FAPE-IR FAPE-IR PR PR-B
De-rainstreak MUSIQ 59.90 60.06 61.02 62.23 59.94 61.00 60.87 62.58 63.06 62.57 62.98 60.23 61.07 59.53 62.62 62.78
Affine-NR -0.90 -0.90 -0.92 -0.91 -0.91 -0.90 -0.92 -0.89 -0.96 -0.96 -0.94 -0.89 -1.00 -1.00 -1.02 -1.03
DR-Score 30.85 30.05 33.77 33.62 51.68 31.48 38.83 32.87 49.07 71.72 25.50 28.37 64.73 76.32 72.12 74.69
Deblur MUSIQ 34.28 32.57 41.64 46.61 43.92 32.85 34.10 49.26 69.99 72.78 55.26 65.73 44.89 45.84 53.37 56.56
Affine-NR -0.73 -0.70 -0.78 -0.77 -0.76 -0.70 -0.72 -0.81 -1.04 -1.10 -0.87 -1.00 -0.80 -0.83 -0.88 -0.93
DR-Score 26.56 31.08 28.12 48.78 45.22 26.58 29.35 46.77 66.66 74.94 38.09 72.18 63.35 65.41 65.43 73.08
De-raindrop MUSIQ 64.18 63.47 64.04 66.26 54.46 61.85 60.62 63.87 65.81 57.74 65.60 66.28 52.71 39.61 47.99 54.54
Affine-NR -0.83 -0.78 -0.82 -0.87 -0.71 -0.77 -0.71 -0.78 -0.83 -0.79 -0.92 -1.02 -0.85 -0.76 -0.72 -0.80
DR-Score 25.70 26.15 22.10 37.63 44.65 26.57 32.38 26.00 34.30 62.59 36.05 36.03 80.73 80.82 74.88 80.62
Desnow MUSIQ 58.96 59.27 59.92 59.74 59.36 60.24 60.15 60.30 63.35 63.46 58.33 62.12 57.96 58.68 62.69 62.65
Affine-NR -0.74 -0.75 -0.77 -0.78 -0.77 -0.75 -0.75 -0.73 -0.83 -0.86 -0.76 -0.89 -0.86 -0.88 -0.90 -0.91
DR-Score 26.90 34.17 39.53 45.99 46.69 28.01 32.45 32.13 46.90 64.89 38.63 50.63 71.57 71.73 71.86 74.70
Dehaze MUSIQ 59.83 60.40 59.66 61.00 60.21 60.26 60.55 60.85 63.38 61.32 63.39 60.32 59.77 60.33 61.93 61.47
Affine-NR -0.88 -0.88 -0.87 -0.88 -0.87 -0.87 -0.88 -0.82 -0.90 -0.86 -0.89 -0.83 -0.88 -0.90 -0.93 -0.93
DR-Score 32.90 37.68 24.55 32.36 28.86 31.02 32.43 47.43 42.61 37.52 40.37 40.74 32.67 35.94 44.51 42.06
Low-light MUSIQ 47.67 55.03 54.21 64.66 53.18 49.61 57.10 48.47 63.33 61.31 54.10 52.75 49.91 57.60 58.17 58.63
Affine-NR -0.88 -0.92 -0.71 -0.95 -0.90 -0.89 -0.98 -0.85 -1.00 -0.94 -0.93 -0.91 -0.89 -0.99 -0.98 -0.99
DR-Score 31.65 67.03 59.33 66.24 65.10 31.70 65.05 38.87 71.25 66.62 58.07 54.71 38.70 75.09 62.38 62.13
Average MUSIQ 54.14 55.13 56.75 60.08 55.18 54.30 55.56 57.55 64.82 63.20 59.45 61.24 54.39 53.60 57.80 59.44
Affine-NR -0.83 -0.82 -0.81 -0.86 -0.82 -0.81 -0.83 -0.81 -0.93 -0.92 -0.89 -0.92 -0.88 -0.88 -0.91 -0.93
DR-Score 29.09 37.69 34.57 44.10 47.03 29.23 38.42 37.51 51.80 63.05 39.45 47.11 58.63 67.55 65.20 67.88
Refer to caption
Figure 7: Visual comparisons on real-world desnow, dehaze, and deblur cases. PixRestore can remove degradations more effectively and recover cleaner details and sharper structures with fewer artifacts.
Refer to caption
Figure 8: Visual comparisons on real-world de-raindrop, low-light enhancement, and de-rainstreak cases. PixRestore can remove degradations more effectively and recover cleaner details and sharper structures with fewer artifacts.

4.6  Generalization to Real-world Test Set

We test all methods on real-world test data to evaluate their generalization ability to out-of-domain scenarios. Since no GT images are available, we report MUSIQ and AFINE-NR as auxiliary no-reference quality metrics, while relying primarily on DR-Score and visual comparisons to assess degradation removal. The results are shown in Table 6. We see that retraining on our training data improves the average DR-Score of many existing methods, but the improvements are not consistent across degradation types. For example, retrained DA-CLIP substantially improves its DR-Score on de-rainstreak, from 33.62 to 51.68, but its score decreases on deblur, from 48.78 to 45.22. This indicates that gains on one degradation may come at the expense of performance drop on another under the unified restoration setting.

It can also be seen that some strong competing methods show leading scores of MUSIQ and Afine-NR on some degradation types. For example, FoundIR-v2 and FoundIR-v2 obtain the highest average MUSIQ scores, but perform poorly on raindrop removal. Flux-IR performs the best on MUSIQ for de-raindrop, but its degradation removal performance on rainstreak is poor. As we discussed in Sec. 4.1, the NR-IQA metrics such as MUSIQ and Afine-NR cannot reflect the real degradation removal performance, and this is why we propose DR-Score to more reliably measure this important ability. From Table 6, we see that PixRestore-B achieves the best average DR-Score of 67.88, followed by FAPE-IR at 67.55. This suggests that PixRestore offers a better overall balance between degradation removal and perceptual quality.

Figs. 7 and 8 present visual comparisons on six real-world degradation types. We see that the original versions of many baseline methods often fail to remove the target degradation sufficiently, while their retrained versions on our training data shown better degradation removal performance, yet they still suffer from residual artifacts, over-smoothing, color bias, or unstable detail reconstruction. In contrast, PixRestore demonstrates more balanced restoration results, achieving stronger degradation removal together with more natural color, clearer structures, and fewer artifacts. For example, in the desnow case, some competing methods leave visible snow residues, whereas PixRestore restores a cleaner image with better structural clarity. In the dehaze example, many competing methods produce grayish or flat-looking outputs with limited visibility, while PixRestore reveals clearer scene content and more natural contrast.

 5  Conclusion

We presented PixRestore, a VAE-free pixel-space diffusion transformer for unified image restoration. By performing flow matching directly on patchified pixels and incorporating adaptive hierarchical DINO guidance in both conditioning and supervision, PixRestore achieved faithful restoration with strong perceptual quality across diverse degradations while remaining compact and efficient with single-step inference. Larger PixRestore variants further improve performance, demonstrating the scalability of our design.

Limitations. PixRestore has several limitations. First, due to the DiT architecture and fixed tokenization setting, a model trained at one resolution cannot be directly extended to higher-resolution without retraining or architectural modification. Second, compared with billion-scale T2I models, the compact PixRestore backbone may encounter difficulties under extremely information-scarce degradations. Finally, our DR-Score relies on a proprietary VLM, introducing potential differences across model updates. In future work, we will explore more resolution-flexible architectures, stronger visual priors, and dedicated IR-specific quality metrics.

References

  • Zhang et al. [2017] Kai Zhang, Wangmeng Zuo, Yunjin Chen, Deyu Meng, and Lei Zhang. Beyond a gaussian denoiser: Residual learning of deep cnn for image denoising. IEEE transactions on image processing, 26(7):3142–3155, 2017.
  • Li et al. [2019] Boyi Li, Wenqi Ren, Dengpan Fu, Dacheng Tao, Dan Feng, Wenjun Zeng, and Zhangyang Wang. Benchmarking single-image dehazing and beyond. IEEE Transactions on Image Processing, 28(1):492–505, 2019.
  • Wang et al. [2024a] Cong Wang, Jinshan Pan, Wei Wang, Gang Fu, Siyuan Liang, Mengzhu Wang, Xiao-Ming Wu, and Jun Liu. Correlation matching transformation transformers for uhd image restoration. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 5336–5344, 2024a.
  • Jiang et al. [2025] Junjun Jiang, Zengyuan Zuo, Gang Wu, Kui Jiang, and Xianming Liu. A survey on all-in-one image restoration: Taxonomy, evaluation and future trends. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.
  • Li et al. [2022a] Boyun Li, Xiao Liu, Peng Hu, Zhongqin Wu, Jiancheng Lv, and Xi Peng. All-In-One Image Restoration for Unknown Corruption. In IEEE Conference on Computer Vision and Pattern Recognition, New Orleans, LA, June 2022a.
  • Potlapalli et al. [2023] Vaishnav Potlapalli, Syed Waqas Zamir, Salman Khan, and Fahad Shahbaz Khan. Promptir: Prompting for all-in-one blind image restoration. Advances in Neural Information Processing Systems (NeurIPS), 2023.
  • Zamfir et al. [2024] Eduard Zamfir, Zongwei Wu, Nancy Mehta, Yuedong Tan, Danda Pani Paudel, Yulun Zhang, and Radu Timofte. Complexity experts are task-discriminative learners for any image restoration, 2024.
  • Luo et al. [2024] Ziwei Luo, Fredrik K Gustafsson, Zheng Zhao, Jens Sjölund, and Thomas B Schön. Photo-realistic image restoration in the wild with controlled vision-language models. arXiv preprint arXiv:2404.09732, 2024.
  • Li et al. [2025a] Hao Li, Xiang Chen, Jiangxin Dong, Jinhui Tang, and Jinshan Pan. Foundir: Unleashing million-scale training data to advance foundation models for image restoration. In Proceedings of the IEEE/CVF international conference on computer vision, pages 12626–12636, 2025a.
  • Zheng et al. [2024] Dian Zheng, Xiao-Ming Wu, Shuzhou Yang, Jian Zhang, Jian-Fang Hu, and Wei-Shi Zheng. Selective hourglass mapping for universal image restoration based on diffusion model. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 25445–25455, 2024.
  • Chen et al. [2025a] Xiang Chen, Jinshan Pan, Jiangxin Dong, Jian Yang, and Jinhui Tang. Foundir-v2: Optimizing pre-training data mixtures for image restoration foundation model. arXiv preprint arXiv:2512.09282, 2025a.
  • Podell et al. [2023] Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952, 2023.
  • Liu et al. [2024a] Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 26296–26306, 2024a.
  • Zhu et al. [2025] Zhiyu Zhu, Jinhui Hou, Hui Liu, Huanqiang Zeng, and Junhui Hou. Learning efficient and effective trajectories for differential equation-based image restoration. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025.
  • Labs [2024] Black Forest Labs. Flux. https://blackforestlabs.ai/announcing-black-forest-labs/, 2024.
  • Peebles and Xie [2023] William Peebles and Saining Xie. Scalable diffusion models with transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4195–4205, 2023.
  • Li and He [2025] Tianhong Li and Kaiming He. Back to basics: Let denoising generative models denoise. URL https://arxiv. org/abs/2511.13720, 7, 2025.
  • Oquab et al. [2024] Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy V. Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel HAZIZA, Francisco Massa, Alaaeldin El-Nouby, Mido Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Michael Rabbat, Vasu Sharma, Gabriel Synnaeve, Hu Xu, Herve Jegou, Julien Mairal, Patrick Labatut, Armand Joulin, and Piotr Bojanowski. DINOv2: Learning robust visual features without supervision. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. URL https://openreview.net/forum?id=a68SUt6zFt. Featured Certification.
  • Zhang and Patel [2018] He Zhang and Vishal M Patel. Density-aware single image de-raining using a multi-stream dense network. In CVPR, 2018.
  • Lin et al. [2024] Jingbo Lin, Zhilu Zhang, Wenbo Li, Renjing Pei, Hang Xu, Hongzhi Zhang, and Wangmeng Zuo. Unirestorer: Universal image restoration via adaptively estimating image degradation at proper granularity. arXiv preprint arXiv:2412.20157, 2024.
  • Lin et al. [2026] Ziyue Lin, Jiahe Hou, Hongyu Xia, Xinrui Xie, Feifei Wang, Yuyin Zhou, Wei Wang, Jiawei Liu, and Liangqiong Qu. Decoupled residual denoising diffusion models for unified and data efficient image-to-image translation. arXiv preprint arXiv:2606.01048, 2026.
  • Ai et al. [2024] Yuang Ai, Huaibo Huang, Xiaoqiang Zhou, Jiexiang Wang, and Ran He. Multimodal prompt perceiver: Empower adaptiveness generalizability and fidelity for all-in-one image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 25432–25444, 2024.
  • Liu et al. [2025] Jingren Liu, Shuning Xu, Qirui Yang, Yun Wang, Xiangyu Chen, and Zhong Ji. Fape-ir: Frequency-aware planning and execution framework for all-in-one image restoration. arXiv preprint arXiv:2511.14099, 2025.
  • Rombach et al. [2022] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022.
  • Ma et al. [2026] Zehong Ma, Ruihan Xu, and Shiliang Zhang. Pixelgen: Pixel diffusion beats latent diffusion with perceptual loss. arXiv preprint arXiv:2602.02493, 2026.
  • Sun et al. [2023] Lingchen Sun, Rongyuan Wu, Zhengqiang Zhang, Hongwei Yong, and Lei Zhang. Improving the stability of diffusion models for content consistent super-resolution. arXiv preprint arXiv:2401.00877, 2023.
  • Yao et al. [2025] Jingfeng Yao, Bin Yang, and Xinggang Wang. Reconstruction vs. generation: Taming optimization dilemma in latent diffusion models. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 15703–15712, 2025.
  • Chen et al. [2023] Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, et al. Pixart-alpha: Fast training of diffusion transformer for photorealistic text-to-image synthesis. arXiv preprint arXiv:2310.00426, 2023.
  • Wu et al. [2025] Chenfei Wu, Jiahao Li, Jingren Zhou, Junyang Lin, Kaiyuan Gao, Kun Yan, Sheng ming Yin, Shuai Bai, Xiao Xu, Yilei Chen, Yuxiang Chen, Zecheng Tang, Zekai Zhang, Zhengyi Wang, An Yang, Bowen Yu, Chen Cheng, Dayiheng Liu, Deqing Li, Hang Zhang, Hao Meng, Hu Wei, Jingyuan Ni, Kai Chen, Kuan Cao, Liang Peng, Lin Qu, Minggang Wu, Peng Wang, Shuting Yu, Tingkun Wen, Wensen Feng, Xiaoxiao Xu, Yi Wang, Yichang Zhang, Yongqiang Zhu, Yujia Wu, Yuxuan Cai, and Zenan Liu. Qwen-image technical report, 2025. URL https://arxiv.org/abs/2508.02324.
  • Radford et al. [2021] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning Transferable Visual Models From Natural Language Supervision. arXiv e-prints, art. arXiv:2103.00020, February 2021.
  • He et al. [2022] Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022.
  • Zhai et al. [2023] Xiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, and Lucas Beyer. Sigmoid loss for language image pre-training. In Proceedings of the IEEE/CVF international conference on computer vision, pages 11975–11986, 2023.
  • Cho et al. [2021] Sung-Jin Cho, Seo-Won Ji, Jun-Pyo Hong, Seung-Won Jung, and Sung-Jea Ko. Rethinking coarse-to-fine approach in single image deblurring. In Proceedings of the IEEE/CVF international conference on computer vision, pages 4641–4650, 2021.
  • Agustsson and Timofte [2017] Eirikur Agustsson and Radu Timofte. Ntire 2017 challenge on single image super-resolution: Dataset and study. In Proceedings of the IEEE conference on computer vision and pattern recognition workshops, pages 126–135, 2017.
  • Xu et al. [2018] J Xu, H Li, Z Liang, D Zhang, and L Zhang. Real-world noisy image denoising: A new benchmark. arXiv preprint arXiv:1804.02603, 2018.
  • Quan et al. [2021] Ruijie Quan, Xin Yu, Yuanzhi Liang, and Yi Yang. Removing raindrops and rain streaks in one go. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 9147–9156, 2021.
  • Li et al. [2022b] Wei Li, Qiming Zhang, Jing Zhang, Zhen Huang, Xinmei Tian, and Dacheng Tao. Toward real-world single image deraining: A new benchmark and beyond. arXiv preprint arXiv:2206.05514, 2022b.
  • Chang et al. [2024] Wenhui Chang, Hongming Chen, Xin He, Xiang Chen, and Liangduo Shen. Uav-rain1k: A benchmark for raindrop removal from uav aerial imagery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15–22, 2024.
  • Li et al. [2023] Chongyi Li, Chun-Le Guo, Man Zhou, Zhexin Liang, Shangchen Zhou, Ruicheng Feng, and Chen Change Loy. Embedding fourier for ultra-high-definition low-light image enhancement. arXiv preprint arXiv:2302.11831, 2023.
  • Wei et al. [2018] Chen Wei, Wenjing Wang, Wenhan Yang, and Jiaying Liu. Deep retinex decomposition for low-light enhancement. arXiv preprint arXiv:1808.04560, 2018.
  • Guan et al. [2025] Qiyuan Guan, Qianfeng Yang, Xiang Chen, Tianyu Song, Guiyue Jin, and Jiyu Jin. Weatherbench: A real-world benchmark dataset for all-in-one adverse weather image restoration. In Proceedings of the 33rd ACM international conference on multimedia, pages 12607–12613, 2025.
  • Cai et al. [2019] Jianrui Cai, Hui Zeng, Hongwei Yong, Zisheng Cao, and Lei Zhang. Toward real-world single image super-resolution: A new benchmark and a new model. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3086–3095, 2019.
  • Wu et al. [2026] Rongyuan Wu, Lingchen Sun, Zhengqiang Zhang, Xiangtao Kong, Jixin Zhao, Shihao Wang, and Lei Zhang. Vosr: A vision-only generative model for image super-resolution. arXiv preprint arXiv:2604.03225, 2026.
  • Lin et al. [2025] Yunlong Lin, Zixu Lin, Haoyu Chen, Panwang Pan, Chenxin Li, Sixiang Chen, Wen Kairun, Yeying Jin, Wenbo Li, and Xinghao Ding. Jarvisir: Elevating autonomous driving perception with intelligent image restoration. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025.
  • Liu et al. [2018] Yun-Fu Liu, Da-Wei Jaw, Shih-Chia Huang, and Jenq-Neng Hwang. Desnownet: Context-aware deep network for snow removal. IEEE Transactions on Image Processing, 27(6):3064–3073, 2018.
  • Ren et al. [2020] Dongwei Ren, Wei Shang, Pengfei Zhu, Qinghua Hu, Deyu Meng, and Wangmeng Zuo. Single image deraining using bilateral recurrent network. IEEE Transactions on Image Processing, 29:6852–6863, 2020.
  • Chen et al. [2025b] I Chen, Wei-Ting Chen, Yu-Wei Liu, Yuan-Chun Chiang, Sy-Yen Kuo, Ming-Hsuan Yang, et al. Unirestore: Unified perceptual and task-oriented image restoration model using diffusion prior. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 17969–17979, 2025b.
  • Wang et al. [2004] Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, 13(4):600–612, 2004.
  • Zhang et al. [2018] Richard Zhang, Phillip Isola, Alexei A Efros, Eli Shechtman, and Oliver Wang. The unreasonable effectiveness of deep features as a perceptual metric. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 586–595, 2018.
  • Ding et al. [2020] Keyan Ding, Kede Ma, Shiqi Wang, and Eero P Simoncelli. Image quality assessment: Unifying structure and texture similarity. IEEE transactions on pattern analysis and machine intelligence, 44(5):2567–2581, 2020.
  • Ke et al. [2021] Junjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar, and Feng Yang. Musiq: Multi-scale image quality transformer. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 5148–5157, 2021.
  • Chen et al. [2025c] Du Chen, Tianhe Wu, Kede Ma, and Lei Zhang. Toward generalized image quality assessment: Relaxing the perfect reference quality assumption. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 12742–12752, 2025c.
  • DeepMind [2026] Google DeepMind. Gemini 3.1 Pro. https://deepmind.google/models/model-cards/gemini-3-1-pro/, 2026.
  • Wang et al. [2024b] Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, et al. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution. arXiv preprint arXiv:2409.12191, 2024b.
  • Kong et al. [2024] Xiangtao Kong, Chao Dong, and Lei Zhang. Towards effective multiple-in-one image restoration: A sequential and prompt learning strategy. arXiv preprint arXiv:2401.03379, 2024.
  • Yang et al. [2026] Yufeng Yang, Xianfang Zeng, Zhangqi Jiang, Fukun Yin, Jianzhuang Liu, Wei Cheng, Shiyu Liu, Yuqi Peng, Gang YU, Shifeng Chen, et al. Realrestorer: Towards generalizable real-world image restoration with large-scale image editing models. arXiv preprint arXiv:2603.25502, 2026.
  • Rim et al. [2020] Jaesung Rim, Haeyun Lee, Jucheol Won, and Sunghyun Cho. Real-world blur dataset for learning and benchmarking deblurring algorithms. In European conference on computer vision, pages 184–201. Springer, 2020.
  • Li et al. [2025b] Yuhao Li, Haoran Fang, Xiang Lei, Qi Wang, Gang Hu, Jiaqing Dong, Zilong Li, Jiabin Lin, Qiegen Liu, and Xianlin Song. Real-world defocus deblurring via score-based diffusion models. Scientific Reports, 15(1):22942, 2025b.
  • Zhang et al. [2025] Jingdong Zhang, Lingzhi Zhang, Qing Liu, Mang Tik Chiu, Connelly Barnes, Yizhou Wang, Haoran You, Xiaoyang Liu, Yuqian Zhou, Zhe Lin, et al. Uniser: A foundation model for unified soft effects removal. arXiv preprint arXiv:2511.14183, 2025.
  • Gómez et al. [2025] Jose L Gómez, Manuel Silva, Antonio Seoane, Agnès Borrás, Mario Noriega, Germán Ros, Jose A Iglesias-Guitian, and Antonio M López. All for one, and one for all: Urbansyn dataset, the third musketeer of synthetic driving scenes. Neurocomputing, 637:130038, 2025.
  • Yao et al. [2020] Yao Yao, Zixin Luo, Shiwei Li, Jingyang Zhang, Yufan Ren, Lei Zhou, Tian Fang, and Long Quan. Blendedmvs: A large-scale dataset for generalized multi-view stereo networks. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 1790–1799, 2020.
  • Li and Snavely [2018] Zhengqi Li and Noah Snavely. Megadepth: Learning single-view depth prediction from internet photos. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 2041–2050, 2018.
  • Jin et al. [2024] Yeying Jin, Xin Li, Jiadong Wang, Yan Zhang, and Malu Zhang. Raindrop clarity: A dual-focused dataset for day and night raindrop removal. In European Conference on Computer Vision, pages 1–17. Springer, 2024.
  • Zamir et al. [2022] Syed Waqas Zamir, Aditya Arora, Salman Khan, Munawar Hayat, Fahad Shahbaz Khan, and Ming-Hsuan Yang. Restormer: Efficient transformer for high-resolution image restoration. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5728–5739, 2022.
  • Zhang et al. [2021] Kaihao Zhang, Rongqing Li, Yanjiang Yu, Wenhan Luo, and Changsheng Li. Deep dense multi-scale network for snow removal using semantic and geometric priors. IEEE Transactions on Image Processing, 2021.
  • Abdelhamed et al. [2018] Abdelrahman Abdelhamed, Stephen Lin, and Michael S Brown. A high-quality denoising dataset for smartphone cameras. In IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 1692–1700, 2018.
  • [67] Flickr. Website. https://www.flickr.com.
  • Yang et al. [2019] Yuanhao Yang, Zheng Luo, Yuhui Chen, et al. Dark face: Face detection in low light condition. CVPR 2019 Workshop on UG2+ Challenge, 2019. URL https://flyywh.github.io/CVPRW2019LowLight/.
  • Liu et al. [2024b] Xiaoning Liu, Zongwei Wu, Ao Li, Florin-Alexandru Vasluianu, Yulun Zhang, Shuhang Gu, Le Zhang, Ce Zhu, Radu Timofte, Zhi Jin, et al. Ntire 2024 challenge on low light image enhancement: Methods and results. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 6571–6594, 2024b.
  • Yang et al. [2024] Heemin Yang, Jaesung Rim, Seungyong Lee, Seung-Hwan Baek, and Sunghyun Cho. Gyro-based neural single image deblurring. arXiv preprint arXiv:2404.00916, 2024.
  • Loh and Chan [2019] Yuen Peng Loh and Chee Seng Chan. Getting to know low-light images with the exclusively dark dataset. Computer Vision and Image Understanding, 178:30–42, 2019. doi: https://doi.org/10.1016/j.cviu.2018.10.010.
  • Chen et al. [2024] Chaofeng Chen, Jiadi Mo, Jingwen Hou, Haoning Wu, Liang Liao, Wenxiu Sun, Qiong Yan, and Weisi Lin. Topiq: A top-down approach from semantics to distortions for image quality assessment. IEEE Transactions on Image Processing, 33:2404–2418, 2024.
  • Yang et al. [2022] Sidi Yang, Tianhe Wu, Shuwei Shi, Shanshan Lao, Yuan Gong, Mingdeng Cao, Jiahao Wang, and Yujiu Yang. Maniqa: Multi-dimension attention network for no-reference image quality assessment. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1191–1200, 2022.

The following materials are provided in this appendix:

  • Details of training data and real-world test data collection (see Sec. 4.1 of the main paper).

  • Pixel-space vs. latent-space diffusion models for unified image restoration (see Sec. 3.1 of the main paper).

  • Visual foundation prior for unified image restoration (see Sec. 3.2 of the main paper).

  • Details of DR-Score (see Sec. 4.1 of the main paper).

  • More public benchmark comparisons, including the per-dataset numerical comparisons and more visual comparisons (Sec. 4.2 of the main paper).

Appendix A Training Data and Real-World Test Data

We build a training corpus of about 2.83M images spanning eight restoration tasks: deblur, dehaze, denoise, de-rainstreak, de-raindrop, desnowing, low-light enhancement, and super-resolution (SR). During training, samples are drawn from these tasks with equal probability. Table S.1 summarizes the training data sources for each degradation type, highlighting the diversity of both degradation patterns and scene content.

For haze, rainstreak, and snow, which depend strongly on scene depth, we further incorporate commonly used datasets with image-depth pairs and synthesize degradations following the pipelines of MioIR [55] and RealRestorer [56]. For denoising, we synthesize noisy inputs by adding Gaussian noise with three noise levels, i.e., σ=15,25, and 50, following DnCNN [1]. For SR, we evaluate the 4× setting, with low-quality (LQ) inputs resized to 512×512 to match the high-quality (HQ) images before feeding them into the model. There is no image-level overlap between the training data and the evaluated public benchmarks. For datasets without an official train/test split, such as PolyU [35] and ScreenSR [43], we reserve a portion of about 10% for testing and use the rest for training.

For real-world evaluation, we collect benchmarks covering six degradation types, each containing 100 real photographs, as summarized in Table S.2. These images are used to evaluate the generalization ability of restoration methods and do not have ground-truth (GT) references. We do not include real-world denoising and SR in this benchmark because they are difficult to define as isolated degradations in real images. In practice, real-world low-quality (LQ) images usually contain mixed degradations: low-light images are often accompanied by noticeable noise, while real-world low-resolution images are also commonly affected by blur and noise. Therefore, we focus on six representative real-world degradation types, which are sufficient to assess the generalization performance of different methods. All real-world images are resized to 512×512 for testing.

媒体内容 · 前往原文查看
Table S.1: Training data sources for each degradation type.
Degradation Training data
Deblur GoPro [33], RealBlur [57], UHD-Blur [3], LSD-Defocus [58]
Dehaze RESIDE [2], UHD-Haze [3], WeatherBench-haze [41], UniSer-Haze [59], synthetic data from UrbanSyn [60], BlendedMVS [61] and MegaDepth [62]
De-raindrop RaindropClarity [63], RainDS-Real-RainDrop [36],
de-rainstreak Rain13K [64], RealRain-1k [37], UAV-Rain1k [38], FoundIR-rain [9], RainDS-Real-RainStreak [36], synthetic data from UrbanSyn [60],
Desnow Snow100K [65], WeatherBench-snow [41], synthetic data from UrbanSyn [60],
Denoise SIDD [66], PolyU [35], synthetic Gaussian noise DF2K [34, 67]
Low-light
enhancement LOL [40], UHD-LL [39], DarkFace [68], FoundIR-low-light [9], NTIRE-LLIE [69]
Super-resolution (SR) RealESRGAN degradation from DF2K [34, 67], RealSR [42], ScreenSR [43]
媒体内容 · 前往原文查看
Table S.2: Real-world test data sources for each degradation type.
Degradation Real-world test data
Deblur GyroBlur-Real [70]
Dehaze RTTS and OpenReal-fog [44]
De-raindrop OpenReal-raindrop [44]
De-rainstreak OpenReal-rainstreak [44] and DiffUIR[10]
Desnow Snow100K-realistic [65] and OpenReal-snow [44]
Low-light
enhancement OpenReal-night [44] and ExDark [71]
媒体内容 · 前往原文查看
Table S.3: Pixel-space and latent-space comparison on 8 degradation types. We report comparisons on PSNR/LPIPS/MUSIQ.
Degradation Latent DiT with FLUX-VAE Latent DiT with Qwen-VAE Latent DiT with SD2-VAE Pixel DiT
SR 25.59/0.2038/58.68 25.62/0.1992/62.25 25.03/0.2586/58.53 26.48/0.2067/57.80
Deblur 26.66/0.1617/46.53 26.93/0.1545/46.69 26.00/0.1921/45.23 27.17/0.1850/41.76
Dehaze 15.31/0.2253/54.56 15.16/0.2391/52.50 15.08/0.2659/53.31 24.78/0.1020/58.17
Denoise 31.97/0.0952/47.22 32.33/0.1021/48.37 30.65/0.1252/47.90 32.58/0.1179/51.32
De-raindrop 20.43/0.2257/62.79 20.63/0.2463/64.71 20.01/0.2751/62.46 23.19/0.1707/66.66
De-rainstreak 25.07/0.1478/46.49 25.34/0.1821/47.63 24.53/0.1876/46.00 30.19/0.1350/49.43
Desnow 27.82/0.1304/49.00 28.41/0.1190/49.39 27.30/0.1483/49.45 28.96/0.1301/48.58
Low-light enhancement 10.82/0.4572/40.64 10.82/0.4530/42.13 10.81/0.4834/39.73 20.80/0.2122/57.98
Overall 22.63/0.2109/50.86 22.80/0.2181/51.87 22.10/0.2483/50.38 26.62/0.1593/54.32

Appendix B Pixel-space vs. Latent-space

In the main paper, we compare pixel-space and latent-space diffusion models using the average results over all eight degradation types, and pixel diffusion shows a clear advantage for unified image restoration (UIR). Here, we further present the per-degradation results in Table S.3.

We conduct the comparison under the same LightningDiT-S [27] backbone and training protocol. The pixel model directly operates on RGB patches with a patch size of 8, while the latent models take latent representations encoded by FLUX-VAE [15], Qwen-VAE [29], and SD2-VAE [24] with a latent patch size of 1. This design keeps the same spatial compression ratio across models for a fair comparison. None of these models uses an external visual foundation model, such as DINOv2 [18]. During inference, all models use 10 sampling steps and a guidance scale of 1.0. We evaluate them on 15 public benchmarks covering eight degradation types. Following the evaluation protocol in Sec. 4.1 of the main paper, we first compute each metric on each test dataset, and then average the results with equal weight within each degradation type.

More specifically, on SR and deblurring, the latent diffusion model with Qwen-VAE achieves better perceptual scores of LPIPS (0.1992 vs. 0.2067 on SR, 0.1545 vs. 0.1850 on deblurring) and MUSIQ (62.25 vs. 57.80 on SR, 46.69 vs. 41.76 on deblurring). For dehazing, de-raindrop removal, de-rainstreak removal, and low-light enhancement, Pixel DiT performs best on all three metrics. The gains are especially large on dehazing (24.78 PSNR, 0.1020 LPIPS, and 58.17 MUSIQ) and low-light enhancement (20.80 PSNR, 0.2122 LPIPS, and 57.98 MUSIQ), far surpassing all latent alternatives. For denoising, Pixel DiT achieves the best PSNR and MUSIQ (32.58 and 51.32), while FLUX-VAE gives the best LPIPS (0.0952). For desnowing, Pixel DiT leads in PSNR (28.96), whereas Qwen-VAE and SD2-VAE lead in LPIPS (0.1190) and MUSIQ (49.45), respectively. Overall, latent models can achieve better perceptual scores on several specific degradations, but pixel-space modeling delivers consistently the best restoration fidelity across diverse tasks.

Appendix C Visual Foundation Prior

We leverage a frozen visual foundation encoder to provide dense visual cues, as the main paper presents. To choose the encoder, we compare four frozen candidates: CLIP [30], DINOv2 [18], MAE [31], and SigLIP [32]. All candidates share the same ViT-B architecture and parameter count, and we extract dense tokens from the same layer position (the 11th layer). Meanwhile, the pixel DiT, training data, and optimization schedule are kept fixed, and only the frozen encoder differs, so that the comparison cleanly reflects the effectiveness of different pretrained vision models for UIR. During inference, all models use 10 sampling steps and a guidance scale of 1.0. Table S.4 reports metrics averaged over 15 public benchmarks covering 8 degradation types. The self-supervised DINOv2 achieves the best overall fidelity and perceptual balance, while CLIP obtains a marginally higher MUSIQ score. We attribute this to its pretraining objective. Through self-distillation over global and local views, DINO learns dense, spatially precise tokens that preserve the fine structures and textures that restoration must recover, while its tokens remain semantically discriminative and encode high-level content. This dual property is what UIR needs: spatial details are used to reconstruct faithful pixels, and semantic cues are used to distinguish reliable content from degradation. In contrast, CLIP and SigLIP are aligned to text and emphasize global semantics, discarding much of the spatial detail, while MAE targets low-level pixel reconstruction and yields less discriminative structural cues. We therefore adopt DINOv2 as the default encoder.

媒体内容 · 前往原文查看
Table S.4: Comparison of frozen visual foundation priors under the same pixel DiT setting. All encoders share the ViT-B architecture and parameter count. Features of the same layer are selected. Metrics are averaged over 8 degradation types. Best results are highlighted in bold.
Visual Prior PSNR (dB)  SSIM  LPIPS  MUSIQ 
CLIP-B 26.61 0.8441 0.1598 54.30
MAE-B 26.70 0.8454 0.1593 54.20
SigLIP-B 26.66 0.8449 0.1592 54.19
DINOv2-B 27.25 0.8500 0.1531 54.21
Refer to caption
Figure S.1: Illustration of DR-Score. Given the LQ input, the restored result, and the task description, the VLM judges whether the target degradation has been removed. A better restoration receives a higher DR-Score.
Refer to caption
Figure S.2: Human alignment analysis of different no-reference metrics with pairwise human preferences. Left: alignment ratios between metric-induced rankings and human judgments across real-world UIR tasks. Right: representative failure cases where two restored results are visually similar with subtle appearance differences but lead to disagreement between DR-Score and human judgment.

Appendix D DR-Score Details

As discussed in the main paper, common full-reference metrics require GT images, which are unavailable for many real-world benchmarks. Existing no-reference metrics mainly assess overall visual quality, but they often fail to capture whether the target degradation has been removed from the restored image. To address this issue, we introduce DR-Score, a vision-language model (VLM)-based auxiliary metric to evaluate degradation removal performance in UIR tasks, rather than to replace standard metrics.

Evaluation Protocol. We use Gemini 3.1 pro [53] as the evaluator, as shown in Fig. S.1. For each test sample, we provide the VLM with the low-quality (LQ) input image, the restored output image, and the restoration task type. Each task type is paired with a short description of the target goal. The VLM compares the restored image with the LQ input and is asked to judge whether the target task has been completed. Specifically, for deblur, the VLM is asked to judge whether the motion or defocus blur has been removed and sharp details are recovered. For dehaze, VLM judges whether the haze or fog has been removed and whether the scene becomes clearer. For de-rainstreak and de-raindrop, it judges whether rain streaks or raindrops are removed and whether the occluded background is restored. For desnow, it judges whether snow is removed. For low-light enhancement, it judges whether the image is properly brightened, denoised, and detailed. For denoising, it judges whether noise is removed while textures are preserved. For SR, it judges whether sharp details are restored without over-smoothing or obvious artifacts. The score ranges from 0 to 100, with 100 indicating that the target degradation is removed completely and cleanly and 0 indicating that it is still fully present.

We also ask the VLM to return the task type, the score, and a short reason. As shown in Fig. S.1, a higher score indicates better task completion, while a lower score indicates that the degradation or restoration artifacts remain. In this way, DR-Score provides a simple auxiliary signal for degradation removal evaluation.

Stability of DR-Score. Because VLM outputs can be stochastic, we evaluate each image five times and report the average score in all experiments. To verify stability, we calculate two complementary statistics.

(i) Run-level standard deviation (std). For each method, we first average scores over all test images within each run and then compute the std across repeated runs. The resulting std value is about 0.10–0.43, indicating that the mean scores at method-level are highly reproducible.

(ii) Image-level std. For each image, we compute the std across its repeated scores and then average over images. The std value is about 5.3–6.4, showing that individual-image scores can vary.

Despite per-image score fluctuations, the std of DR-Scores at method-level is below 0.5, confirming that the relative ranking between methods is stable. We therefore report mean DR-scores in the experiments.

Human Alignment. We then conduct a human study to examine how well different metrics agree with human judgment in UIR tasks, including DR-Score, MUSIQ [51], AFINE-NR [52], TOPIQ [72], and MANIQA [73].

The study is based on pairwise comparisons. For each test case, annotators are shown one LQ image and two restored results produced by different methods for the same task. They are asked to select the better result of the restoration task. The annotators are instructed to focus on three aspects: whether the target degradation is removed, whether the main scene content is preserved, and whether obvious artifacts are suppressed. We sample 20 cases for each degradation type from the real-world test sets, including deblur, dehaze, de-rainstreak, de-raindrop, desnow, and low-light enhancement. For each pair, two methods are randomly selected from retrained models (DA-CLIP, FoundIR, FoundIR-v2, FAPE-IR) and PixRestore. The two results are presented in random order to avoid position bias. We also hide the method names and do not show any metric values.

A total of 20 annotators participated in this study. After collecting the annotations, we average the human choices and convert them into pairwise preferences. We then compare these preferences with the pairwise ranking induced by each metric. A pair is counted as aligned if the image preferred by humans also receives a better metric score. We report the average alignment ratio in the left part of Fig. S.2. DR-Score achieves the highest agreement (90.7%) with human preference, showing that it better reflects task completion and degradation removal than existing no-reference quality metrics. This supports its use as a practical auxiliary metric for real-world UIR evaluation.

Failure Cases. DR-Score still has several limitations. It can be less reliable when two restored results are visually very similar and differ only in subtle appearance factors, such as brightness, local contrast, or color tone, as shown in the right part of Fig. S.2. In such cases, even human annotators need to inspect carefully, and the VLM may fail to tell the difference.

媒体内容 · 前往原文查看
Table S.5: Detailed quantitative comparison on Deblur benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with are retrained under the same training setting as ours.
Method GoPro UHD-blur
PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR
PromptIR 23.43 0.7903 0.3035 0.1972 25.50 -0.5874 24.21 0.7265 0.3228 0.2418 29.81 -0.6163
PromptIR 29.82 0.8775 0.2000 0.1440 36.18 -0.7123 28.39 0.8172 0.2101 0.1681 40.52 -0.7570
DiffUIR 29.32 0.8669 0.2041 0.1492 35.93 -0.7137 26.39 0.7731 0.2476 0.1907 38.01 -0.7075
UniRestore 24.11 0.7496 0.2216 0.1408 45.93 -0.7896 23.63 0.6914 0.2595 0.2024 42.99 -0.6870
DA-CLIP 28.57 0.8554 0.1279 0.0994 41.05 -0.7329 26.23 0.7671 0.2003 0.1612 42.83 -0.7479
DA-CLIP 28.87 0.8536 0.1317 0.1031 42.91 -0.7444 27.71 0.7920 0.1630 0.1264 47.05 -0.7656
FoundIR 27.02 0.8121 0.2610 0.1800 29.96 -0.6276 27.56 0.7971 0.2250 0.1762 38.80 -0.7544
FoundIR 29.95 0.8762 0.1887 0.1393 35.81 -0.7297 28.33 0.8170 0.2084 0.1613 39.98 -0.7696
FoundIR-v2 24.48 0.7179 0.2224 0.1451 57.73 -0.8954 24.32 0.6980 0.2183 0.1680 63.23 -1.0422
FoundIR-v2 24.98 0.7470 0.1965 0.1262 56.90 -0.8914 24.97 0.7232 0.1862 0.1385 58.39 -0.9774
Flux-IR 24.03 0.7012 0.2408 0.1550 55.90 -0.8500 22.98 0.6479 0.3084 0.2153 53.29 -0.8064
Flux-IR 22.59 0.6716 0.2717 0.1936 60.22 -0.9609 21.71 0.6176 0.2918 0.2137 58.22 -0.9313
FAPE-IR 28.02 0.8374 0.1527 0.1058 41.62 -0.7781 25.58 0.7417 0.2669 0.2054 35.00 -0.7267
FAPE-IR 28.68 0.8505 0.1347 0.0907 38.84 -0.7403 27.25 0.7908 0.1807 0.1282 41.57 -0.7694
PixRestore-S 28.96 0.8553 0.1067 0.0833 43.39 -0.8182 27.67 0.8016 0.1335 0.1046 49.73 -0.8105
PixRestore-B 30.00 0.8805 0.0895 0.0721 44.74 -0.8354 28.46 0.8218 0.1206 0.0949 49.54 -0.8200
PixRestore-L 30.79 0.8948 0.0787 0.0657 45.51 -0.8616 28.96 0.8346 0.1117 0.0891 50.45 -0.8515
PixRestore-XL 31.23 0.9025 0.0721 0.0621 46.34 -0.8712 29.07 0.8407 0.1090 0.0854 50.34 -0.8487
媒体内容 · 前往原文查看
Table S.6: Detailed quantitative comparison on Dehaze benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with are retrained under the same training setting as ours.
Method RESIDE-6K UHD-Haze
PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR
PromptIR 26.69 0.9572 0.0460 0.0418 51.24 -0.9422 15.98 0.8035 0.2391 0.1605 63.94 -0.9765
PromptIR 24.01 0.9417 0.0642 0.0553 51.58 -0.9362 18.49 0.8445 0.2054 0.1216 63.45 -1.0058
DiffUIR 24.66 0.9311 0.0708 0.0582 50.59 -0.9314 16.15 0.8001 0.2572 0.1768 62.57 -0.9683
UniRestore 23.63 0.9162 0.1171 0.0869 55.68 -0.9630 16.59 0.7776 0.3073 0.1849 62.14 -0.9155
DA-CLIP 28.66 0.9273 0.0559 0.0463 53.52 -0.9784 17.27 0.8230 0.2063 0.1307 65.88 -1.0698
DA-CLIP 28.15 0.9579 0.0431 0.0409 51.31 -0.9439 15.70 0.7979 0.2487 0.1746 63.98 -0.9540
FoundIR 16.52 0.8381 0.1665 0.1304 49.68 -0.8612 13.62 0.7421 0.3499 0.2533 59.75 -0.8267
FoundIR 25.07 0.9525 0.0523 0.0483 50.61 -0.9464 16.91 0.8282 0.2146 0.1441 64.37 -0.9976
FoundIR-v2 19.01 0.8101 0.1837 0.1318 57.83 -0.9396 19.11 0.7309 0.2018 0.1357 69.04 -1.0056
FoundIR-v2 18.57 0.7925 0.1997 0.1342 56.12 -0.9191 19.62 0.7365 0.1870 0.1222 68.08 -1.0543
Flux-IR 15.64 0.7796 0.2591 0.1639 65.46 -1.0185 13.83 0.7402 0.3361 0.2333 61.74 -0.8161
Flux-IR 17.27 0.8416 0.1578 0.1074 50.54 -0.8348 14.65 0.7706 0.3052 0.2083 60.48 -0.8251
FAPE-IR 31.36 0.9628 0.0378 0.0364 50.35 -0.9573 19.20 0.8386 0.1542 0.0937 66.12 -1.1555
FAPE-IR 28.90 0.9575 0.0411 0.0385 50.32 -0.9488 21.78 0.8537 0.1442 0.0905 65.45 -1.1361
PixRestore-S 28.41 0.9555 0.0473 0.0444 52.81 -0.9777 22.51 0.8729 0.1318 0.0855 66.98 -1.1836
PixRestore-B 29.87 0.9636 0.0390 0.0388 51.79 -0.9737 23.41 0.8820 0.1188 0.0760 67.26 -1.2451
PixRestore-L 30.63 0.9660 0.0365 0.0366 52.11 -0.9772 23.23 0.8836 0.1227 0.0771 66.81 -1.2381
PixRestore-XL 31.73 0.9677 0.0340 0.0355 51.83 -0.9741 23.14 0.8827 0.1223 0.0776 66.92 -1.2237
媒体内容 · 前往原文查看
Table S.7: Detailed quantitative comparison on Denoise benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with are retrained under the same training setting as ours.
Method DIV2K (Gaussian) PolyU
PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR
PromptIR 34.57 0.9079 0.1411 0.1384 64.24 -0.9098 30.50 0.8978 0.3506 0.1825 30.88 -0.6068
PromptIR 33.91 0.8975 0.1463 0.1388 61.00 -0.8517 37.00 0.9780 0.0744 0.0742 32.70 -0.6156
DiffUIR 21.24 0.7540 0.3437 0.2523 55.21 -0.7013 31.98 0.9221 0.2769 0.1588 32.67 -0.6270
UniRestore 30.38 0.8728 0.1659 0.1420 64.25 -0.8868 33.11 0.9209 0.3012 0.1918 31.60 -0.5629
DA-CLIP 28.91 0.7532 0.2412 0.1764 57.33 -0.8155 25.79 0.8957 0.2506 0.1765 33.00 -0.6137
DA-CLIP 30.76 0.7704 0.2448 0.1570 58.12 -0.7929 37.19 0.9753 0.0477 0.0761 33.67 -0.5932
FoundIR 27.14 0.6061 0.4920 0.2574 45.98 -0.6097 37.77 0.9789 0.0668 0.0707 33.55 -0.6303
FoundIR 34.06 0.8992 0.1572 0.1469 62.82 -0.9074 38.25 0.9837 0.0529 0.1022 34.24 -0.6195
FoundIR-v2 25.43 0.6544 0.2658 0.1825 62.76 -0.9197 28.31 0.8418 0.2913 0.2066 52.61 -0.8119
FoundIR-v2 25.96 0.6584 0.2488 0.1837 62.60 -0.9080 30.39 0.8410 0.3290 0.1943 41.44 -0.6478
Flux-IR 21.33 0.4780 0.5758 0.2660 50.37 -0.7322 30.08 0.8903 0.2637 0.2073 43.26 -0.7832
Flux-IR 24.34 0.5837 0.4296 0.2386 49.22 -0.7338 29.26 0.8984 0.3281 0.1716 32.85 -0.6348
FAPE-IR 31.09 0.8540 0.1147 0.1013 60.49 -0.8739 34.89 0.9636 0.1329 0.1212 35.29 -0.6526
FAPE-IR 31.32 0.8572 0.1080 0.0978 60.52 -0.8830 37.11 0.9772 0.0401 0.0522 33.52 -0.6135
PixRestore-S 33.04 0.8923 0.0783 0.0863 63.62 -0.9410 36.71 0.9749 0.0465 0.0607 34.54 -0.6139
PixRestore-B 33.40 0.8968 0.0725 0.0777 63.84 -0.9591 35.84 0.9745 0.0402 0.0636 34.15 -0.6030
PixRestore-L 33.35 0.8962 0.0730 0.0768 64.12 -0.9671 35.91 0.9736 0.0385 0.0709 34.39 -0.6022
PixRestore-XL 33.46 0.8979 0.0729 0.0774 64.06 -0.9648 36.50 0.9745 0.0341 0.0753 33.95 -0.5994
媒体内容 · 前往原文查看
Table S.8: Detailed quantitative comparison on De-rainstreak benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with are retrained under the same training setting as ours.
Method RainDS-real RealRain-1K
PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR
PromptIR 25.15 0.7483 0.2047 0.1392 60.54 -0.8602 23.85 0.7531 0.5011 0.3370 42.82 -0.6698
PromptIR 26.62 0.7933 0.1977 0.1254 61.79 -0.8963 30.23 0.8845 0.3341 0.2554 36.49 -0.5757
DiffUIR 26.11 0.7888 0.1885 0.1168 63.88 -0.9157 22.91 0.7348 0.5084 0.3329 45.21 -0.7440
UniRestore 23.47 0.7143 0.3220 0.1838 65.71 -0.8797 21.52 0.7439 0.5226 0.3540 45.97 -0.7192
DA-CLIP 24.66 0.7467 0.1834 0.1213 63.20 -0.9041 24.35 0.7654 0.4861 0.3139 46.39 -0.7354
DA-CLIP 25.72 0.7516 0.1432 0.0859 63.31 -0.9261 37.50 0.9695 0.0665 0.0910 32.27 -0.6034
FoundIR 26.78 0.7890 0.1630 0.1075 63.26 -0.9045 26.97 0.8635 0.3278 0.2523 39.15 -0.6486
FoundIR 27.27 0.8120 0.1993 0.1191 67.29 -0.9812 37.44 0.9655 0.1193 0.1216 35.47 -0.6340
FoundIR-v2 23.65 0.6078 0.1956 0.1111 63.58 -0.9722 22.69 0.7226 0.5362 0.3326 48.23 -0.7375
FoundIR-v2 23.74 0.6043 0.1906 0.1088 64.82 -1.0001 31.96 0.9153 0.1641 0.1556 36.17 -0.6297
Flux-IR 22.34 0.6665 0.2695 0.1769 64.41 -0.9676 19.63 0.5801 0.6554 0.3952 50.43 -0.7374
Flux-IR 21.66 0.6410 0.2912 0.1803 62.00 -0.9306 20.37 0.6278 0.6110 0.3840 45.93 -0.6811
FAPE-IR 26.53 0.7636 0.1407 0.0835 62.03 -0.9748 28.50 0.8817 0.3232 0.2523 42.51 -0.7340
FAPE-IR 26.67 0.7642 0.1270 0.0746 61.14 -0.9816 37.14 0.9748 0.0535 0.0778 31.15 -0.6092
PixRestore-S 27.06 0.7950 0.1180 0.0773 66.07 -1.0433 37.50 0.9744 0.0625 0.1038 32.07 -0.6152
PixRestore-B 27.41 0.8038 0.1078 0.0693 65.65 -1.0461 38.28 0.9758 0.0456 0.0941 32.27 -0.6153
PixRestore-L 27.54 0.8065 0.0990 0.0664 65.60 -1.0532 38.51 0.9760 0.0404 0.0933 32.65 -0.6170
PixRestore-XL 27.58 0.8079 0.1000 0.0650 65.18 -1.0463 38.60 0.9753 0.0401 0.1028 32.81 -0.6160
媒体内容 · 前往原文查看
Table S.9: Detailed quantitative comparison on De-raindrop benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with are retrained under the same training setting as ours.
Method RainDS-real UAV-Rain1k
PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR
PromptIR 20.70 0.7069 0.2773 0.1560 55.31 -0.8088 16.92 0.6855 0.4007 0.2211 66.57 -0.8022
PromptIR 24.53 0.7529 0.2637 0.1347 59.14 -0.8560 22.84 0.8482 0.1835 0.1305 68.19 -0.8475
DiffUIR 20.56 0.6996 0.3119 0.1678 56.28 -0.8034 17.13 0.7011 0.3790 0.2132 66.67 -0.7942
UniRestore 20.25 0.6906 0.3482 0.1888 58.44 -0.8111 16.75 0.5847 0.4746 0.2683 65.96 -0.7044
DA-CLIP 22.99 0.7018 0.1763 0.0977 61.66 -0.8832 17.26 0.6958 0.3667 0.2047 67.27 -0.8080
DA-CLIP 24.14 0.7077 0.1570 0.0905 60.92 -0.8855 23.01 0.8755 0.1079 0.0799 69.43 -0.9228
FoundIR 20.67 0.7171 0.3145 0.1749 59.34 -0.8041 17.07 0.6773 0.4026 0.2285 67.57 -0.7823
FoundIR 25.41 0.7708 0.2487 0.1319 65.20 -0.9316 23.40 0.8729 0.1397 0.0973 69.58 -0.9111
FoundIR-v2 20.22 0.5725 0.3079 0.1566 59.84 -0.8704 19.03 0.5194 0.2505 0.1560 69.47 -0.8975
FoundIR-v2 21.72 0.5641 0.2435 0.1255 64.75 -0.9876 19.91 0.5280 0.2165 0.1332 69.30 -0.9315
Flux-IR 21.01 0.6540 0.2335 0.1221 63.08 -1.0046 16.87 0.6377 0.4113 0.2328 66.83 -0.8472
Flux-IR 18.85 0.5672 0.3313 0.1964 68.66 -1.1729 16.76 0.5897 0.3218 0.2186 72.29 -1.0842
FAPE-IR 24.71 0.7255 0.1841 0.0930 59.93 -0.9452 18.06 0.6389 0.2598 0.1702 65.88 -0.8070
FAPE-IR 25.27 0.7263 0.1502 0.0790 58.40 -0.9358 22.49 0.7183 0.1607 0.1145 68.35 -0.9177
PixRestore-S 25.42 0.7394 0.1326 0.0761 61.54 -0.9911 23.54 0.8117 0.1190 0.1002 70.15 -0.9816
PixRestore-B 25.76 0.7461 0.1245 0.0720 61.41 -0.9928 24.65 0.8490 0.0928 0.0806 70.35 -1.0176
PixRestore-L 25.89 0.7504 0.1195 0.0686 61.30 -0.9963 25.16 0.8631 0.0810 0.0719 70.63 -1.0379
PixRestore-XL 26.01 0.7567 0.1171 0.0677 61.36 -0.9939 25.69 0.8774 0.0715 0.0655 70.55 -1.0392
媒体内容 · 前往原文查看
Table S.10: Detailed quantitative comparison on Low-light Enhancement benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with are retrained under the same training setting as ours.
Method UHD-LL LOL
PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR
PromptIR 11.82 0.5660 0.5088 0.3116 35.06 -0.6577 9.17 0.3902 0.5732 0.4415 38.49 -0.7859
PromptIR 25.33 0.8846 0.2303 0.1757 45.85 -0.6794 10.50 0.4916 0.4824 0.3268 44.53 -0.8607
DiffUIR 17.67 0.5064 0.6057 0.3320 37.19 -0.4779 25.76 0.9099 0.1581 0.1264 68.79 -0.9938
UniRestore 12.41 0.6102 0.4665 0.2903 36.98 -0.6417 9.48 0.4274 0.5347 0.3426 43.85 -0.7569
DA-CLIP 20.51 0.7436 0.3675 0.2201 48.51 -0.6172 23.99 0.8395 0.1263 0.1043 74.18 -1.0050
DA-CLIP 16.64 0.7776 0.2632 0.2016 47.98 -0.7589 19.38 0.8594 0.1626 0.1212 66.85 -0.9450
FoundIR 14.22 0.6984 0.3548 0.2386 42.38 -0.7037 16.47 0.7963 0.2519 0.1877 65.83 -1.0337
FoundIR 24.28 0.8885 0.2113 0.1687 51.48 -0.8552 22.40 0.9168 0.1471 0.1191 71.35 -1.0455
FoundIR-v2 16.40 0.7120 0.3643 0.2381 60.71 -0.9313 17.95 0.7776 0.2620 0.1674 67.72 -1.0373
FoundIR-v2 17.72 0.7272 0.3541 0.2273 55.85 -0.8751 16.36 0.7341 0.2834 0.1929 62.16 -0.9620
Flux-IR 15.35 0.5519 0.5455 0.2965 38.47 -0.5140 22.36 0.8525 0.1658 0.1032 72.10 -1.1667
Flux-IR 14.33 0.6028 0.4866 0.2852 37.25 -0.5275 22.47 0.8369 0.1906 0.1167 70.91 -1.2086
FAPE-IR 12.57 0.6315 0.3899 0.2669 39.07 -0.7489 26.34 0.8942 0.1261 0.1006 66.74 -1.0265
FAPE-IR 26.55 0.8980 0.1445 0.1096 51.54 -0.8156 25.53 0.9049 0.1296 0.1040 65.50 -1.0535
PixRestore-S 26.28 0.8889 0.1420 0.1060 53.31 -0.8461 24.96 0.8899 0.1300 0.0957 66.07 -1.0386
PixRestore-B 26.36 0.8886 0.1384 0.1010 53.45 -0.8534 25.08 0.8983 0.1212 0.0899 65.88 -1.0329
PixRestore-L 26.89 0.8928 0.1326 0.0979 54.53 -0.8613 25.30 0.8960 0.1273 0.0932 64.39 -1.0055
PixRestore-XL 26.37 0.8934 0.1324 0.0982 54.62 -0.8686 26.20 0.8955 0.1286 0.0949 63.64 -1.0015
媒体内容 · 前往原文查看
Table S.11: Detailed quantitative comparison on Desnow benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with are retrained under the same training setting as ours.
Method WeatherBench
PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR
PromptIR 22.26 0.7939 0.2452 0.1672 45.60 -0.6023
PromptIR 29.32 0.8532 0.1818 0.1402 46.30 -0.6161
DiffUIR 22.95 0.7948 0.2392 0.1667 48.10 -0.6185
UniRestore 22.33 0.7863 0.2464 0.1771 49.39 -0.6451
DA-CLIP 23.60 0.7971 0.2221 0.1558 46.48 -0.6127
DA-CLIP 28.31 0.8282 0.1360 0.1021 48.82 -0.6424
FoundIR 23.03 0.7999 0.2406 0.1630 45.77 -0.6029
FoundIR 29.82 0.8678 0.1524 0.1224 47.98 -0.6908
FoundIR-v2 24.72 0.7347 0.2513 0.1671 59.98 -0.8032
FoundIR-v2 26.28 0.7690 0.1824 0.1311 54.44 -0.7440
Flux-IR 21.74 0.7231 0.3434 0.2204 56.08 -0.6879
Flux-IR 21.79 0.6831 0.3023 0.1959 57.03 -0.7504
FAPE-IR 26.02 0.8191 0.1759 0.1189 46.29 -0.6280
FAPE-IR 30.19 0.8676 0.1136 0.0849 47.84 -0.6595
PixRestore-S 31.26 0.8859 0.0853 0.0669 49.86 -0.6829
PixRestore-B 31.93 0.8959 0.0688 0.0584 50.10 -0.6825
PixRestore-L 32.35 0.9039 0.0656 0.0569 50.36 -0.6859
PixRestore-XL 32.57 0.9077 0.0623 0.0553 50.25 -0.6864
媒体内容 · 前往原文查看
Table S.12: Detailed quantitative comparison on Super-resolution benchmarks. The best and second-best results are highlighted in red and blue italic, respectively. Methods marked with are retrained under the same training setting as ours.
Method RealSR ScreenSR
PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR PSNR SSIM LPIPS DISTS MUSIQ AFINE-NR
PromptIR 23.47 0.7380 0.4647 0.2670 25.95 -0.4435 25.05 0.7365 0.4140 0.2351 42.95 -0.7419
PromptIR 28.65 0.8051 0.3045 0.2382 45.99 -0.6821 26.54 0.7884 0.2633 0.2033 58.41 -0.8518
DiffUIR 27.40 0.7734 0.3744 0.2417 35.15 -0.4999 25.68 0.7431 0.3989 0.2338 48.34 -0.7720
UniRestore 24.84 0.7683 0.3358 0.2305 39.91 -0.5857 24.76 0.7403 0.3737 0.2263 53.30 -0.8091
DA-CLIP 24.21 0.7258 0.3910 0.2469 30.60 -0.4804 23.24 0.6321 0.3635 0.2265 47.89 -0.7682
DA-CLIP 27.72 0.7777 0.2213 0.1821 50.30 -0.7260 25.26 0.7393 0.2469 0.1705 64.42 -0.9523
FoundIR 26.13 0.7432 0.4474 0.2619 26.74 -0.4396 25.58 0.7365 0.4086 0.2365 42.63 -0.7437
FoundIR 28.74 0.8000 0.3354 0.2454 40.54 -0.6443 26.56 0.7903 0.2412 0.2048 61.52 -0.9192
FoundIR-v2 24.61 0.6682 0.3303 0.2248 68.03 -1.0134 23.09 0.6640 0.2616 0.1700 65.93 -0.9700
FoundIR-v2 24.83 0.6649 0.3206 0.2179 64.84 -0.9785 22.74 0.6411 0.1781 0.1185 72.46 -1.0921
Flux-IR 23.42 0.6599 0.3548 0.2465 69.74 -1.0971 21.56 0.6483 0.2258 0.1664 72.17 -1.2027
Flux-IR 22.41 0.5927 0.3668 0.2559 68.62 -1.0386 19.34 0.5778 0.3471 0.2479 71.41 -1.1189
FAPE-IR 27.92 0.7969 0.2325 0.1877 50.67 -0.8189 25.17 0.7495 0.3307 0.2077 52.83 -0.8243
FAPE-IR 29.06 0.8139 0.1843 0.1468 50.37 -0.8224 25.90 0.7546 0.2061 0.1398 61.45 -0.8721
PixRestore-S 28.40 0.7882 0.1776 0.1458 56.20 -0.8501 25.62 0.7578 0.1696 0.1258 66.40 -0.9682
PixRestore-B 28.43 0.7904 0.1646 0.1367 56.15 -0.8590 25.71 0.7642 0.1556 0.1211 67.20 -0.9940
PixRestore-L 28.55 0.7933 0.1590 0.1317 57.26 -0.8880 25.46 0.7508 0.1483 0.1227 68.72 -0.9846
PixRestore-XL 28.59 0.7956 0.1595 0.1332 56.37 -0.8922 25.17 0.7399 0.1504 0.1353 68.07 -0.9560
Refer to caption
Figure S.3: Visual comparisons on synthetic desnow, low-light enhancement, de-rainstreak, and denoise cases. Overall, PixRestore restores cleaner and more faithful results with better structural details and fewer artifacts.
Refer to caption
Figure S.4: Visual comparisons on synthetic dehaze, de-raindrop, deblur, and SR cases. Overall, PixRestore restores cleaner and more faithful results with better structural details and fewer artifacts.

Appendix E More Public Benchmark Comparisons

In the main paper, we report the average results of each degradation type. In this appendix, we further provide detailed comparisons on each benchmark, including GoPro [33] and UHD-Blur [3] for deblurring, RESIDE-6K [2] and UHD-Haze [3] for dehazing, DIV2K [34] with Gaussian noise and PolyU [35] for denoising, RainDS-real [36] and RealRain-1k [37] for rain streak removal, RainDS-real [36] and UAV-Rain1k [38] for raindrop removal, UHD-LL [39] and LOL [40] for low-light enhancement, WeatherBench [41] for desnowing, and RealSR [42] and ScreenSR [43] for super-resolution. All images are center-cropped to 512 for testing.

The results are shown in Tables S.5S.12. We see that PixRestore is not limited to a specific test dataset. It achieves strong and balanced performance across diverse restoration benchmarks. Some previous methods can obtain good no-reference scores by producing sharper or more contrastive outputs, but they often fall behind on fidelity-oriented full-reference metrics. For example, in the deblur task, FoundIR-v2 and Flux-IR obtain much higher MUSIQ and better AFINE-NR on GoPro and UHD-Blur, but their PSNR, SSIM, LPIPS, and DISTS are clearly worse than PixRestore. Compared with previous UIR methods, our model consistently ranks among the top methods on distortion metrics such as PSNR, SSIM, LPIPS, and DISTS. At the same time, it remains competitive on no-reference metrics such as MUSIQ and AFINE-NR.

We can see that there is a clear trend across almost all tasks: scaling the model size is beneficial. From PixRestore-S to PixRestore-B, to PixRestore-L, and to PixRestore-XL, scaling generally improves average performance, although some individual datasets show non-monotonic behavior. This trend is especially clear for deblur, dehaze, desnow, and de-raindrop, where larger models repeatedly deliver stronger restoration fidelity. Although the gains on MUSIQ or AFINE-NR are sometimes less monotonic, the larger variants still show more stable top-tier performance overall. These results suggest that PixRestore scales well, and that increasing model capacity is an effective way to improve UIR performance.

The results of retrained models using our training data also reveal an important pattern. Many existing methods retrained on our training data improve their performance on almost all datasets. Nonetheless, these retrained baselines remain behind PixRestore on the main full-reference metrics. This suggests that using better training data alone is not sufficient; the model design itself also matters.

We provide more visual comparisons in Figs. S.3 and S.4. Overall, the compared methods show different trade-offs between degradation removal and detail preservation. PixRestore consistently produces cleaner and more balanced results across diverse synthetic tasks. For example, in the desnow case of Fig. S.3, PixRestore removes snow more thoroughly while preserving fine fence structures. Similar trends can be observed in Fig. S.4. In deblurring, PixRestore restores sharper pole boundaries and cleaner background tree textures; in super-resolution, it recovers clearer brick patterns and more faithful structures than competing methods.

来源:HuggingFace Daily Papers(社区热门论文)· arxiv.org