晨光
暗夜
晨光
极光
Bilingual Paper Reading · 中英对照精读

RoadFusion:面向路面缺陷检测的潜扩散模型

准大一 · 土木工程 × 道路工程 × 计算机视觉 × AI —— 路面缺陷检测精读材料
原文:arXiv:2507.15346 2025年7月21日发布 arXiv 预印本(cs.CV) 路面缺陷检测 × 潜扩散模型 × 双适配器 附英文摘要朗读音频

一、论文档案

英文标题RoadFusion: Latent Diffusion Model for Pavement Defect Detection
中文标题RoadFusion:面向路面缺陷检测的潜扩散模型
作者穆罕默德·阿基尔, 基杜斯·达格纳·贝莱特, 弗朗切斯科·塞蒂(机构未在素材中标注)
发布时间2025年7月21日(v1)|分类:cs.CV(计算机视觉)
一句话概括标注数据不够、场景千变万化,就用「潜扩散模型凭空生成逼真病害图」来补数据,再让两个特征适配器分别处理正常/异常输入——在 6 个路面基准数据集上刷新多项纪录。
💡 为什么选这篇给你:① 道路工程是土木最「接地气」的方向之一,路面病害(裂缝、坑槽、磨损)检测直接服务养护决策;② 思路清晰好懂:用「生成模型造数据」化解「缺陷样本太少」的痛点,大一也能读懂主线;③ 六个公开基准数据集验证 + 分类与定位双任务,故事完整、可复现。

二、核心术语表(先扫一遍再读正文)

英文术语中文大白话解释
pavement defect detection路面缺陷检测用图像自动找出路面上的裂缝、坑槽、磨损等病害。
latent diffusion model潜扩散模型先在压缩的「潜空间」里做扩散生成,再解码成图像——生成又快又逼真。
synthetic anomaly generation合成异常生成用模型「人造」逼真的缺陷图,扩充训练数据。
text prompt文本提示用一句话描述想生成的缺陷(如「细长裂缝」),指导生成方向。
spatial mask空间掩码指定缺陷应出现在图像哪个位置的黑白图。
dual-adaptor architecture双适配器架构两条独立通路,分别适配「正常路面」与「异常路面」的特征。
domain shift域偏移(领域漂移)训练数据和实际数据「长得不一样」(换相机、换天气),模型性能就掉。
feature adaptor特征适配器把预训练模型的通用特征「翻译」成路面专用的特征表示。
discriminator判别器分辨「真假缺陷」的小网络,逼着生成器做得更真。
patch level图像块级别把图像切成小块逐个判断,兼顾细节与效率。
pixel-level localization像素级定位逐像素标出缺陷区域——精细到「每个像素属于缺陷还是背景」。
image-level classification图像级分类只判断整张图「有没有缺陷」,不画位置。
state-of-the-art (SOTA)最先进水平同类任务里目前最好的成绩。
benchmark dataset基准数据集公认的测试标准数据集,所有方法在同一批数据上比高低。

三、摘要中英对照(精读核心)

🎧 音频在文末,可先听一遍原文再读;每个英文句都配了逐句翻译。

摘要 Abstract

EN · 原文
Pavement defect detection faces critical challenges including limited annotated data, domain shift between training and deployment environments, and high variability in defect appearances across different road conditions.
CN · 翻译
路面缺陷检测面临严峻挑战,包括标注数据有限、训练环境与部署环境之间的域偏移(domain shift),以及不同路况下缺陷外观的高变异性
EN · 原文
We propose RoadFusion, a framework that addresses these limitations through synthetic anomaly generation with dual-path feature adaptation.
CN · 翻译
我们提出 RoadFusion——一个通过合成异常生成 + 双通路特征适配来解决上述局限的框架。
EN · 原文
A latent diffusion model synthesizes diverse, realistic defects using text prompts and spatial masks, enabling effective training under data scarcity.
CN · 翻译
潜扩散模型利用文本提示与空间掩码合成多样、逼真的缺陷,使数据稀缺条件下也能有效训练。
EN · 原文
Two separate feature adaptors specialize representations for normal and anomalous inputs, improving robustness to domain shift and defect variability.
CN · 翻译
两个独立的特征适配器分别专精「正常输入」与「异常输入」的表示,提升模型对域偏移与缺陷变异的鲁棒性。
EN · 原文
A lightweight discriminator learns to distinguish fine-grained defect patterns at the patch level.
CN · 翻译
一个轻量判别器学习在图像块(patch)级别区分细粒度的缺陷模式。
EN · 原文
Evaluated on six benchmark datasets, RoadFusion achieves consistently strong performance across both classification and localization tasks, setting new state-of-the-art in multiple metrics relevant to real-world road inspection.
CN · 翻译
六个基准数据集上的评估表明,RoadFusion 在分类与定位任务上均表现稳定,在多个与真实道路巡检相关的指标上刷新最先进水平(SOTA)

关键词 Keywords:Pavement Defect Detection 路面缺陷检测 | Latent Diffusion Model 潜扩散模型 | Synthetic Anomaly Generation 合成异常生成 | Domain Adaptation 域适配

四、引言精选(为什么这个问题重要)

① 道路养护为什么重要:40% 的支出比例很震撼

EN · 原文
Road infrastructure is a cornerstone of national development, underpinning mobility, economic activity, public safety, and territorial accessibility. The structural condition of pavement surfaces directly impacts vehicle performance, fuel efficiency, travel time reliability, and user safety. Poorly maintained roads lead to increased wear on vehicles and higher operating costs. In Europe, road maintenance alone can represent up to 40% of total transport infrastructure spending, emphasizing the importance of targeted and timely maintenance efforts.
CN · 翻译
道路基础设施是国家发展的基石,支撑着出行、经济活动、公共安全与区域可达性。路面结构的状况直接影响车辆性能、燃油效率、出行时间可靠性与用户安全。养护不善的路会加剧车辆磨损、推高运营成本。在欧洲,仅道路养护一项就可能占交通基础设施总支出的 40%——足见「及时、精准养护」的重要性。

② 人工巡检的困境与意大利的例子

EN · 原文
For public administrations (PAs) managing extensive and aging road networks, early detection of surface anomalies is critical. Traditional visual inspection methods are labor-intensive, inconsistent, and lack scalability. In Italy, for example, more than 250,000 kilometers of roads require regular assessment. The national road agency, ANAS, allocates over €1.5 billion annually to pavement rehabilitation, yet resource constraints continue to limit large-scale, proactive maintenance. As a result, AI-driven inspection systems are gaining traction as cost-effective, scalable alternatives.
CN · 翻译
对管理着庞大且日益老化路网的公共管理部门(PA)而言,尽早发现表面异常至关重要。传统人工目视巡检劳动密集、标准不一、难以规模化。以意大利为例,超过 25 万公里道路需要定期评估;国家道路机构 ANAS 每年拨款超过 15 亿欧元用于路面修复,但资源约束仍限制着大规模、主动式养护。因此,AI 驱动的巡检系统正作为高性价比、可扩展的替代方案兴起。

③ 深度学习强在分类,弱在定位

EN · 原文
Recent advances in deep learning-particularly convolutional neural networks (CNNs)-have demonstrated strong performance in tasks such as crack classification, pothole detection, and texture anomaly recognition. However, a key limitation of existing approaches is their primary focus on image-level classification, rather than on precise localization of defects. For real-world deployment, especially in public infrastructure management, simply knowing that a defect exists is insufficient. High-resolution, pixel-level localization is essential for prioritizing maintenance, estimating damage extent, and planning repairs efficiently.
CN · 翻译
近年来深度学习——尤其是卷积神经网络(CNN)——在裂缝分类、坑槽检测、纹理异常识别等任务上表现强劲。但现有方法的关键局限是:主要做整图级分类,而不是缺陷的精确定位。对公共基础设施管理而言,光知道「有缺陷」远远不够——高分辨率的像素级定位才是排定养护优先级、估算损坏程度、高效安排修复的前提。

④ 三大实战挑战:域偏移、类别不平衡、外观多样

EN · 原文
Additionally, several practical challenges persist. First, pre-trained models often struggle with domain shifts when applied to real pavement data. Second, the imbalance between abundant normal samples and limited defect samples reduces training effectiveness. Third, the visual diversity of pavement conditions-across materials, lighting, weather, and imaging perspectives-adds noise and complexity to feature learning. These issues contribute to false positives, missed detections, and unreliable predictions-outcomes that are costly and dangerous in practice.
CN · 翻译
此外还有几个实际挑战:第一,预训练模型应用到真实路面数据时经常遭遇域偏移;第二,正常样本充足、缺陷样本稀少的类别不平衡削弱训练效果;第三,路面外观的多样性——材料、光照、天气、拍摄角度各异——给特征学习增加噪声与复杂度。这些问题导致误检、漏检和不可靠的预测——在实践中既昂贵又危险。
💡 这是全文最值得体会的一段“For real-world deployment, especially in public infrastructure management, simply knowing that a defect exists is insufficient.”——道路巡检的第一性需求不是「知道有没有病」,而是「病在哪、多严重」。问题定义对了,方法才有意义。

五、论文贡献(3 个要点)

EN · 原文
A dual-adaptor architecture that bridges the domain gap between pre-trained features and pavement-specific representations, using separate pathways for normal and anomalous samples to improve discriminative power.
CN · 翻译
1. 双适配器架构。在预训练特征与路面专用表示之间搭建桥梁,为正常与异常样本设置独立通路,提升判别能力。
EN · 原文
Integration of a latent diffusion model for generating diverse, realistic synthetic anomalies guided by text prompts and spatial masks, helping address the scarcity of annotated defect data.
CN · 翻译
2. 潜扩散数据增强。引入潜扩散模型,以文本提示 + 空间掩码生成多样、逼真的合成异常,缓解标注缺陷数据稀缺问题。
EN · 原文
A streamlined inference pipeline that maintains computational efficiency while delivering high-resolution anomaly localization across challenging, real-world datasets.
CN · 翻译
3. 精简推理流程。在保持计算效率的同时,于困难的真实数据集上实现高分辨率异常定位。

六、结论中英对照

EN · 原文
We introduced RoadFusion, a diffusion-based framework for pavement defect detection that combines anomaly generation with dual-adaptor feature learning.
CN · 翻译
本文提出 RoadFusion——一个把「异常生成」与「双适配器特征学习」相结合的扩散式路面缺陷检测框架。
EN · 原文
By leveraging latent diffusion to synthesize diverse pavement anomalies, our approach addresses the limited availability of labeled defect data. The dual feature adaptors enable domain-specific feature alignment, improving defect localization and classification.
CN · 翻译
通过潜扩散合成多样路面异常,解决标注缺陷数据不足的问题;双特征适配器实现领域专属的特征对齐,改善缺陷定位与分类。
EN · 原文
Experiments across six benchmark datasets demonstrate that RoadFusion consistently outperforms existing methods in both detection and localization tasks. Results confirm strong generalization across various road surfaces and defect types, while qualitative samples validate the realism of synthesized anomalies. RoadFusion provides an effective solution for pavement monitoring when annotated data are limited or diverse defect types are expected.
CN · 翻译
六个基准数据集上的实验表明,RoadFusion 在检测与定位任务上均稳定优于现有方法;结果证实了其在多种路面与缺陷类型上的泛化能力,定性样本也验证了合成异常的逼真度。当标注数据有限或缺陷类型多样时,RoadFusion 为路面监测提供了有效方案。

七、编者解读:这篇论文到底讲了什么(大白话版)

  1. 问题:现实里「好路面」的照片一抓一大把,「坏路面」(裂缝、坑槽)的标注图却很少——训练数据严重失衡;而且不同地区的路面长得不一样,模型换个环境就不灵了。
  2. 做法:先用潜扩散模型「造假数据」——告诉它「生成一条细裂缝、放在这个位置」,造出大量逼真缺陷图补进训练集;再用两个特征适配器分别学习「正常路面」和「异常路面」的长相;最后用轻量判别器在图像块级别做精细的缺陷判别。
  3. 结果:在 Cracks and Potholes、Pothole600、Crack500、Edmcrack600、Gaps384、CNR Road 六个公开数据集上,分类与定位任务都取得稳定领先,多项指标刷新纪录。
  4. 最值钱的观点:与其到处收集标注数据,不如「造数据」——生成模型把数据稀缺问题从源头化解;同时「正常/异常分开建模」让模型各有专攻,比一个黑箱更稳。
  5. 工程意义:对养护单位来说,知道「裂缝在哪、多严重」才能排优先级、估预算——这篇论文把「知道有没有」升级成「知道在哪」,离真实落地更近一步。
🎯 对保研的启示:这篇论文示范了「数据痛点 → 生成式解法」的科研叙事:先讲清楚数据稀缺为什么致命,再讲生成模型怎么补,最后用多数据集证明泛化。复试时能讲清「痛点—思路—验证」闭环,比堆砌模型名词更能打动导师。

八、给准大一的阅读路线图 & 延伸方向

📖 怎么读这篇论文(三遍法)

  1. 第一遍(10 分钟):只读摘要和术语表,回答三个问题——问题是什么?方法是什么?结果是什么?
  2. 第二遍(20 分钟):读引言 + 结论,重点体会「为什么要用生成模型造数据」和「像素级定位为什么是刚需」。
  3. 第三遍(30 分钟):读方法文字部分(潜扩散生成、双适配器、判别器),跳过所有公式和编号,只看文字描述;遇到不懂的术语回查术语表。

🚀 这个方向你能延伸做什么

九、英文摘要朗读(练听力用)

先盲听一遍→再看对照稿→再听一遍。目标是听出每个数字(six benchmark datasets、40%、€1.5 billion)和术语(latent diffusion、dual-adaptor、discriminator、domain shift)。