主页 文献库文献详情
PMID: 41807484 已发表 · epublish 英语

Integrated framework utilizing scene text detection and recognition techniques for enhancing point of interest extraction from name boards in all Indic languages.

Scientific reports ·第 16 卷 ·第 1 期 ·2026-03-10

Kashyap AK, Upadhya M, Panwar VS, Chandrakar V

摘要

This paper focused on enhancing text recognition, script identification, language classification, and point of interest (POI) extraction from images captured by Mobile Mapping Systems (MMS). The initiative was undertaken to improve the existing Computer vision-based artificial intelligence modules. The advancements made is to be contributed to the system's implementation, bringing improved functionality and accuracy to the system. The current system, called Text Detection and Recognition (TDR), consists of several neural modules operating in sequential stages. The first module identifies areas of interest within the MMS images, focusing on shop signboards, traffic signs, and directional boards. The second stage involves detecting text words within these areas and cropping the relevant pixels from the image. These cropped images are then processed in the third stage, where the language script is detected, identifying one of ten Indian scripts. In the fourth stage, specific character recognizers corresponding to the identified script are used to recognize the text. The outputs from these stages are aggregated into a correlated JSON output. Additionally, a parallel fifth stage detects various fields within the MMS images, such as name, address, pin, icon, phone and GSTIN number, ultimately extracting a comprehensive human-readable address for any POI from the MMS image. The primary focus includes investigating novel text recognition algorithms to improve accuracy and efficiency, exploring various script identification algorithms to enhance language classification capabilities, implementing a dictionary-based approach for more accurate word detection, and developing methods for correcting the words that the CRNN model predicts to reduce errors. This work is novel because it combines word correction, OCR, detection, classification, and POI field extraction into a single pipeline designed specifically for Indic scripts. By obtaining 96.17% script recognition accuracy, 92.5% word accuracy, and 33% average precision in POI detection, the suggested framework outperforms previous benchmarks like IndicText (93.6%) and transformer-based OCR (88.5%).

关键词
Deep learning-based OCR Multilingual script identification Object detection Point of interest (POI) extraction Scene text recognition
文献信息
期刊
Scientific reports
期刊简称
Sci Rep
ISSN
2045-2322
发表日期
2026-03-10
语言
英语
国家/地区
England
NLM ID
101563288
分析服务
分析服务

联系地址

山东省济南市章丘区文博路2号

齐鲁师范学院 genelibs生信实验室

山东省济南市高新区舜华路750号

大学科技园北区F座4单元2楼

电话: 0531-88819269

微信公众号

关注微信订阅号,实时查看信息,关注医学生物学动态。


商务邮箱

E-mail: [email protected]