Home LiteratureArticle Details
PMID: 40003124 Published · epublish English Journal Article

CrackCLIP: Adapting Vision-Language Models for Weakly Supervised Crack Segmentation.

Entropy (Basel, Switzerland) ·Vol. 27 ·No. 2 ·2025-01-25

Liang F, Li Q, Yu H, Wang W

Abstract

Weakly supervised crack segmentation aims to create pixel-level crack masks with minimal human annotation, which often only differentiate between crack and normal no-crack patches. This task is crucial for assessing structural integrity and safety in real-world industrial applications, where manually labeling the location of cracks at the pixel level is both labor-intensive and impractical. Addressing the challenges of labeling uncertainty, this paper presents CrackCLIP, a novel approach that leverages language prompts to augment the semantic context and employs the Contrastive Language-Image Pre-Training (CLIP) model to enhance weakly supervised crack segmentation. Initially, a gradient-based class activation map is used to generate pixel-level coarse pseudo-labels from a trained crack patch classifier. The estimated coarse pseudo-labels are utilized to fine-tune additional linear adapters, which are integrated into the frozen image encoders of CLIP to adapt the CLIP model to the specialized task of crack segmentation. Moreover, specific textual prompts are crafted for crack characteristics, which are input into the frozen text encoder of CLIP to extract features encapsulating the semantic essence of the cracks. The final crack segmentation is determined by comparing the similarity between text prompt features and visual patch token features. Comparative experiments on the Crack500, CFD, and DeepCrack datasets demonstrate that the proposed framework outperforms existing weakly supervised crack segmentation methods, and the pre-trained vision-language model exhibits strong potential for crack feature learning, thereby enhancing the overall performance and generalization capabilities of the proposed framework.

Keywords
Contrastive Language–Image Pre-Training vision-language model weakly supervised crack segmentation
作者与单位
共 4 位作者,点击展开单位 / ORCID
Liang Fengjiao ORCID
Key Laboratory of Big Data & Artificial Intelligence in Transportation (Beijing Jiaotong University), Ministry of Education, Beijing 100044, China.
Li Qingyong ORCID
Key Laboratory of Big Data & Artificial Intelligence in Transportation (Beijing Jiaotong University), Ministry of Education, Beijing 100044, China. | Frontiers Science Center for Smart High-Speed Railway System, Beijing Jiaotong University, Beijing 100044, China.
Yu Haomin
Department of Computer Sicence, Aalborg University, 9200 Aalborg, Denmark.
Wang Wen
Key Laboratory of Big Data & Artificial Intelligence in Transportation (Beijing Jiaotong University), Ministry of Education, Beijing 100044, China.
Article Info
Journal
Entropy (Basel, Switzerland)
Abbr.
Entropy (Basel)
ISSN
1099-4300
Published
2025-01-25
电子出版
2025-00-25
Language
English
Country/Region
Switzerland
NLM ID
101243874
基金资助
the Fundamental Research Funds for the Central Universities under Grant · 2022JBMC055, 2023JBZY037
the Beijing Natural Science Foundation under Grant · L231019
the Shanghai Industrial Development Project under Grant · HCXBCY-2023-033
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]