Dual-stage gradient-guided diffusion model for image super-resolution
- Pages: 1-11(2026)
Received:28 May 2026,
Revised:2026-07-24,
Accepted:10 August 2026,
Online First:12 August 2026
DOI: 10.11834/jig.260307
移动端阅览

浏览全部资源
扫码关注微信
Received:28 May 2026,
Revised:2026-07-24,
Accepted:10 August 2026,
Online First:12 August 2026,
移动端阅览
目的
2
现有基于卷积网络或Transformer的图像超分辨率方法受限于像素级损失函数的设计,重建结果常出现细节模糊、高频纹理重建不足等问题。近年来扩散模型凭借强大的生成先验在感知质量上取得了显著提升,但在纹理与边缘等细节区域容易生成与输入视觉内容不一致的幻觉纹理。针对上述问题,提出一种双阶段梯度引导的扩散模型图像超分辨率方法,从训练与推理两个阶段同时增强生成过程的视觉一致性。
方法
2
训练阶段引入频域视觉一致性损失,通过前向视觉算子刻画高分辨率图像的频谱结构,并将其梯度注入噪声预测,引导模型更准确地重建高频细节;推理阶段基于扩散后验采样,设计测量一致性与流形粘合的双重约束,对每一步潜变量更新进行校正,使采样轨迹既符合观测退化模型,又保持在自然图像分布内。
结果
2
在RealSR(real-world super-resolution)和DRealSR(diverse real-world super-resolution)两个数据集上与现有主流方法进行对比实验。客观指标方面,与基线扩散超分辨率方法相比,所提方法在RealSR数据集上的峰值信噪比(peak signal-to-noise ratio,PSNR)提升1.756 dB,结构相似性(structural similarity index measure,SSIM)提升0.0115,感知图像块相似度(learned perceptual image patch similarity,LPIPS)降低2.76%;在DRealSR数据集上PSNR提升2.139 dB,SSIM提升0.0203,LPIPS降低10.88%。主观评价方面,该方法在字符纹理等高频细节区域有效抑制了幻觉伪影,结构准确性与纹理真实感优于对比方法。消融实验验证了各模块的有效性及协同增益。
结论
2
所提出的双阶段梯度引导的扩散模型图像超分辨率方法通过训练阶段的频域视觉引导与推理阶段的双重约束后验采样,有效抑制了幻觉纹理,提升了重建结果的保真度与结构一致性,在感知质量与重建保真度的综合权衡上优于现有方法。
Objective
2
Image super-resolution (ISR) aims to reconstruct high-resolution (HR) images from degraded low-resolution (LR) inputs, and is an important task in computer vision and image processing. Conventional methods based on convolutional neural networks (CNNs) or Transformer architectures have achieved considerable progress, but they are usually optimized with pixel-wise losses, such as mean squared error (MSE) or mean absolute error (MAE). These losses tend to produce over-smoothed results and insufficient high-frequency details, especially in regions with fine textures, characters, and sharp edges. In recent years, diffusion models have shown strong generative priors for image super-resolution and can produce perceptually realistic textures. However, their stochastic sampling process may generate hallucinated textures that are inconsistent with the input LR observation, which weakens the structural reliability of reconstructed images. To address this problem, a dual-stage gradient-guided diffusion model for image super-resolution is proposed to enhance visual consistency from both the training and inference stages.
Method
2
The proposed method introduces explicit gradient guidance into two stages of a diffusion-based super-resolution pipeline. In the training stage, a frequency-domain visual consistency loss is used to supervise the noise prediction network. A forward visual operator extracts the spectral representation of the HR target image, and the gradient derived from this frequency-domain loss is injected into the noise prediction process. This strategy encourages the denoising network to learn latent representations that are more consistent with the frequency distribution of real HR images, thereby improving the reconstruction of high-frequency structures without modifying the base network architecture. In the inference stage, a dual-constraint diffusion posterior sampling strategy is designed to correct the latent variable update at each denoising step. The first constraint is measurement consistency, which enforces consistency between the generated HR result and the input LR observation under the degradation model. The second constraint is manifold gluing, which encourages the corrected latent variables to remain close to the natural image manifold and reduces structural drift during sampling. By combining training-stage frequency-domain visual guidance with inference-stage dual-constraint posterior sampling, the proposed framework suppresses hallucinated textures while preserving perceptual quality and structural consistency.
Result
2
Experiments are conducted on two super-resolution benchmarks, RealSR (real-world super-resolution) and DRealSR (diverse real-world super-resolution), under the ×4 upscaling setting. The proposed method is compared with representative super-resolution methods through quantitative and qualitative evaluations. The evaluation metrics include peak signal-to-noise ratio (PSNR) and structural similarity index measure (SSIM) for fidelity assessment, learned perceptual image patch similarity (LPIPS) for perceptual similarity evaluation, CLIP image quality assessment (CLIP-IQA) and multi-scale image quality transformer (MUSIQ) for no-reference perceptual quality evaluation, and natural image quality evaluator (NIQE) for blind image quality assessment. Compared with the baseline diffusion-based super-resolution method, the proposed method improves PSNR by 1.756 dB, improves SSIM by 0.0115, and reduces LPIPS by 2.76% on the RealSR dataset. On the DRealSR dataset, the proposed method improves PSNR by 2.139 dB, improves SSIM by 0.0203, and reduces LPIPS by 10.88%. These results indicate that the proposed method improves reconstruction fidelity on both datasets while maintaining a reasonable balance between fidelity and perceptual quality. Qualitative comparisons further show that the proposed method effectively suppresses hallucination artifacts in high-frequency regions, such as character textures and complex patterns, and generates results with improved structural accuracy and texture fidelity. Ablation experiments verify that both the training-stage frequency-domain guidance and the inference-stage dual-constraint posterior sampling contribute to the final performance, and their combination provides complementary gains.
Conclusion
2
A dual-stage gradient-guided diffusion model for image super-resolution is proposed. By incorporating frequency-domain visual guidance during training and dual-constraint posterior sampling during inference, the proposed framework suppresses hallucinated textures and improves the fidelity and structural consistency of reconstructed images. Experimental results on RealSR and DRealSR demonstrate that the proposed method achieves a better balance between reconstruction fidelity and perceptual quality than the baseline diffusion-based method. Future work will focus on reducing the computational cost of gradient-guided sampling and extending the framework to blind super-resolution with unknown degradations.
Cai J R , Zeng H , Yong H W , Cao Z S and Zhang L . 2019 . Toward real-world single image super-resolution: a new benchmark and a new model // Proceedings of the IEEE/CVF International Conference on Computer Vision . Seoul, Korea (South) : IEEE: 3086 - 3095 [ DOI: 10.1109/ICCV.2019.00318 http://dx.doi.org/10.1109/ICCV.2019.00318 ]
Chen J Y , Pan J S and Dong J X . 2025 . FaithDiff: unleashing diffusion priors for faithful image super-resolution // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Nashville, TN, USA : IEEE: 28188 - 28197 [ DOI: 10.1109/CVPR52734.2025.02625 http://dx.doi.org/10.1109/CVPR52734.2025.02625 ]
Chung H , Kim J , McCann M T , Klasky M L and Ye J C . 2023 . Diffusion posterior sampling for general noisy inverse problems // Proceedings of the 11th International Conference on Learning Representations . Kigali, Rwanda : OpenReview
Efron B . 2011 . Tweedie’s formula and selection bias . Journal of the American Statistical Association , 106 ( 496 ): 1602 - 1614 [ DOI: 10.1198/jasa.2011.tm11181 http://dx.doi.org/10.1198/jasa.2011.tm11181 ]
Ho J , Saharia C , Chan W , Fleet D J , Norouzi M and Salimans T . 2022 . Cascaded diffusion models for high fidelity image generation . Journal of Machine Learning Research , 23 ( 47 ): 1 - 33
Ledig C , Theis L , Huszar F , Caballero J , Cunningham A , Acosta A , et al . 2017 . Photo-realistic single image super-resolution using a generative adversarial network // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition . Honolulu, HI, USA : IEEE Computer Society: 105 - 114 [ DOI: 10.1109/CVPR.2017.19 http://dx.doi.org/10.1109/CVPR.2017.19 ]
Li X Y , Wang Z R , Zou Y , Chen Z X , Ma J , Jiang Z Y , et al . 2025 . DifIISR: a diffusion model with gradient guidance for infrared image super-resolution // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . Nashville, TN, USA : IEEE: 7534 - 7544 [ DOI: 10.1109/CVPR52734.2025.00706 http://dx.doi.org/10.1109/CVPR52734.2025.00706 ]
Li Y C , Liang D , Ding T Y and Huang S J . 2025 . StructSR: refuse spurious details in real-world image super-resolution // Proceedings of the AAAI Conference on Artificial Intelligence . Philadelphia, PA, USA : AAAI Press: 5022 - 5030 [ DOI: 10.1609/aaai.v39i5.32532 http://dx.doi.org/10.1609/aaai.v39i5.32532 ]
Liang J Y , Cao J Z , Sun G L , Zhang K , Van Gool L and Timofte R . 2021 . SwinIR: image restoration using Swin Transformer // Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops . Montreal, QC, Canada : IEEE: 1833 - 1844 [ DOI: 10.1109/ICCVW54120.2021.00210 http://dx.doi.org/10.1109/ICCVW54120.2021.00210 ]
Lim B , Son S H , Kim H W , Nah S J and Lee K M . 2017 . Enhanced deep residual networks for single image super-resolution // Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops . Honolulu, HI, USA : IEEE Computer Society: 1132 - 1140 [ DOI: 10.1109/CVPRW.2017.151 http://dx.doi.org/10.1109/CVPRW.2017.151 ]
Liu Z Y , Yang Y , Huang S Y and Wang S Z . 2026 . Recursive super-resolution network with adaptive feature fusion . Journal of Image and Graphics , 31 ( 4 ): 1044 - 1060
刘紫阳 , 杨勇 , 黄淑英 , 王书昭 . 2026 . 自适应特征融合的超分辨率重建循环网络 . 中国图象图形学报 , 31 ( 4 ): 1044 - 1060 [ DOI: 10.11834/jig.250332 http://dx.doi.org/10.11834/jig.250332 ]
Rombach R , Blattmann A , Lorenz D , Esser P and Ommer B . 2022 . High-resolution image synthesis with latent diffusion models // Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition . New Orleans, LA, USA : IEEE: 10684 - 10695 [ DOI: 10.1109/CVPR52688.2022.01042 http://dx.doi.org/10.1109/CVPR52688.2022.01042 ]
Rout L , Raoof N , Daras G , Caramanis C , Dimakis A and Shakkottai S . 2023 . Solving linear inverse problems provably via posterior sampling with latent diffusion models // Proceedings of the Advances in Neural Information Processing Systems 36 . New Orleans, LA, USA : Curran Associates, Inc.: 49960 - 49990 [ DOI: 10.52202/075280-2174 http://dx.doi.org/10.52202/075280-2174 ]
Saharia C , Ho J , Chan W , Salimans T , Fleet D J and Norouzi M . 2023 . Image super-resolution via iterative refinement . IEEE Transactions on Pattern Analysis and Machine Intelligence , 45 ( 4 ): 4713 - 4726 [ DOI: 10.1109/TPAMI.2022.3204461 http://dx.doi.org/10.1109/TPAMI.2022.3204461 ]
Tan M K , Xu S K , Zhang S H and Chen Q . 2021 . A review on deep adversarial visual generation . Journal of Image and Graphics , 26 ( 12 ): 2751 - 2766
谭明奎 , 许守恺 , 张书海 , 陈奇 . 2021 . 深度对抗视觉生成综述 . 中国图象图形学报 , 26 ( 12 ): 2751 - 2766 [ DOI: 10.11834/jig.210252 http://dx.doi.org/10.11834/jig.210252 ]
Wan Y H , Jiang P T , Hou Q B , Zhang H , Chen J W , Cheng M M , et al . 2024 . ControlSR: taming diffusion models for consistent real-world image super resolution [EB/OL]. [ 2026-05-13 ]. https://arxiv.org/abs/2410.14279v2 https://arxiv.org/abs/2410.14279v2
Wang J Y , Yue Z S , Zhou S C , Chan K C K and Loy C C . 2024 . Exploiting diffusion prior for real-world image super-resolution . International Journal of Computer Vision , 132 ( 12 ): 5929 - 5949 [ DOI: 10.1007/s11263-024-02168-7 http://dx.doi.org/10.1007/s11263-024-02168-7 ]
Wang X D , Wang P , Li Z Y and Yuan X . 2025 . Plug-and-play diffusion models for inverse problems with data consistency projection // Proceedings of the IEEE International Conference on Image Processing Workshops . Anchorage, AK, USA : IEEE: 350 - 355 [ DOI: 10.1109/ICIPW68931.2025.11386346 http://dx.doi.org/10.1109/ICIPW68931.2025.11386346 ]
Wang X T , Xie L B , Dong C and Shan Y . 2021 . Real-ESRGAN: training real-world blind super-resolution with pure synthetic data // Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops . Montreal, QC, Canada : IEEE: 1905 - 1914 [ DOI: 10.1109/ICCVW54120.2021.00217 http://dx.doi.org/10.1109/ICCVW54120.2021.00217 ]
Wei P X , Xie Z W , Lu H N , Zhan Z Y , Ye Q X , Zuo W M , et al . 2020 . Component divide-and-conquer for real-world image super-resolution // Proceedings of the European Conference on Computer Vision . Glasgow, UK : Springer: 101 - 117 [ DOI: 10.1007/978-3-030-58598-3_7 http://dx.doi.org/10.1007/978-3-030-58598-3_7 ]
Wei Y Y , Mao T Y , Li B A , Wang F , Li F , Zhang Z , et al . 2025 . Visual and large multimodal models promote image restoration and enhancement: research progress . Journal of Image and Graphics , 30 ( 5 ): 1197 - 1219
韦炎炎 , 毛天一 , 李柏昂 , 王飞 , 李锋 , 张召 , 等 . 2025 . 视觉模型及多模态大模型推进图像复原增强研究进展. 中国图象图形学报 , 30 ( 5 ): 1197 - 1219 [ DOI: 10.11834/jig.240436 http://dx.doi.org/10.11834/jig.240436 ]
相关文章
相关作者
相关机构
京公网安备11010802024621