Home LiteratureArticle Details
PMID: 24578357 Published · ppublish English Journal Article Research Support, Non-U.S. Gov't

Learning regular expressions for clinical text classification.

Journal of the American Medical Informatics Association : JAMIA ·Vol. 21 ·No. 5 ·2014-00-00 ·Pages 850-7

Bui DD, Zeng-Treitler Q

Abstract

Natural language processing (NLP) applications typically use regular expressions that have been developed manually by human experts. Our goal is to automate both the creation and utilization of regular expressions in text classification. We designed a novel regular expression discovery (RED) algorithm and implemented two text classifiers based on RED. The RED+ALIGN classifier combines RED with an alignment algorithm, and RED+SVM combines RED with a support vector machine (SVM) classifier. Two clinical datasets were used for testing and evaluation: the SMOKE dataset, containing 1091 text snippets describing smoking status; and the PAIN dataset, containing 702 snippets describing pain status. We performed 10-fold cross-validation to calculate accuracy, precision, recall, and F-measure metrics. In the evaluation, an SVM classifier was trained as the control. The two RED classifiers achieved 80.9-83.0% in overall accuracy on the two datasets, which is 1.3-3% higher than SVM's accuracy (p<0.001). Similarly, small but consistent improvements have been observed in precision, recall, and F-measure when RED classifiers are compared with SVM alone. More significantly, RED+ALIGN correctly classified many instances that were misclassified by the SVM classifier (8.1-10.3% of the total instances and 43.8-53.0% of SVM's misclassifications). Machine-generated regular expressions can be effectively used in clinical text classification. The regular expression-based classifier can be combined with other classifiers, like SVM, to improve classification performance.

Keywords
Machine Learning Natural Language Processing Regular Expressions Support Vector Machines Text Classification
MeSH Terms
Algorithms Artificial Intelligence Electronic Data Processing Humans Medical Records Systems, Computerized/classification Natural Language Processing Pain/classification Smoking Support Vector Machine
Authors & Affiliations
2 authors, click to expand affiliations / ORCID
Bui Duy Duc An
Department of Biomedical Informatics, University of Utah, Salt Lake City, Utah, USA VA Salt Lake City Health Care System, Salt Lake City, Utah, USA.
Zeng-Treitler Qing
Department of Biomedical Informatics, University of Utah, Salt Lake City, Utah, USA VA Salt Lake City Health Care System, Salt Lake City, Utah, USA.
References (26)
26 references, click to expand
  1. Combining Lexico-semantic Features for Emotion Classification in Suicide Notes.
    Biomed Inform Insights. 2012;5(Suppl. 1):125-8 PMID: 22879768
  2. A simple algorithm for identifying negated findings and diseases in discharge summaries.
    J Biomed Inform. 2001 Oct;34(5):301-10 PMID: 12123149
  3. PASTE: patient-centered SMS text tagging in a medication management system.
    J Am Med Inform Assoc. 2012 May-Jun;19(3):368-74 PMID: 21984605
  4. Using ensemble models to classify the sentiment expressed in suicide notes.
    Biomed Inform Insights. 2012;5(Suppl. 1):77-85 PMID: 22879763
  5. Use of semantic features to classify patient smoking status.
    AMIA Annu Symp Proc. 2008 Nov 06;:450-4 PMID: 18998969
  6. Identifying data sharing in biomedical literature.
    AMIA Annu Symp Proc. 2008 Nov 06;:596-600 PMID: 18998887
  7. Identifying patient smoking status from medical discharge records.
    J Am Med Inform Assoc. 2008 Jan-Feb;15(1):14-24 PMID: 17947624
  8. A novel hybrid approach to automated negation detection in clinical radiology reports.
    J Am Med Inform Assoc. 2007 May-Jun;14(3):304-11 PMID: 17329723
  9. Active learning for clinical text classification: is it better than random sampling?
    J Am Med Inform Assoc. 2012 Sep-Oct;19(5):809-16 PMID: 22707743
  10. Automatic topic identification of health-related messages in online health community using text classification.
    Springerplus. 2013 Jul 10;2:309 PMID: 23961389
  11. Detection of blood culture bacterial contamination using natural language processing.
    AMIA Annu Symp Proc. 2009 Nov 14;2009:411-5 PMID: 20351890
  12. A hybrid approach to sentiment sentence classification in suicide notes.
    Biomed Inform Insights. 2012;5(Suppl. 1):43-50 PMID: 22879759
  13. Assessing the difficulty and time cost of de-identification in clinical narratives.
    Methods Inf Med. 2006;45(3):246-52 PMID: 16685332
  14. Automated categorisation of clinical incident reports using statistical text classification.
    Qual Saf Health Care. 2010 Dec;19(6):e55 PMID: 20724392
  15. Text Categorization of Heart, Lung, and Blood Studies in the Database of Genotypes and Phenotypes (dbGaP) Utilizing n-grams and Metadata Features.
    Biomed Inform Insights. 2013 Jul 22;6:35-45 PMID: 23926434
  16. Clinical decision support with automated text processing for cervical cancer screening.
    J Am Med Inform Assoc. 2012 Sep-Oct;19(5):833-9 PMID: 22542812
  17. A multi-classifier based guideline sentence classification system.
    Healthc Inform Res. 2011 Dec;17(4):224-31 PMID: 22259724
  18. Automated information extraction of key trial design elements from clinical trial publications.
    AMIA Annu Symp Proc. 2008 Nov 06;:141-5 PMID: 18999067
  19. Deafness mutation mining using regular expression based pattern matching.
    BMC Med Inform Decis Mak. 2007 Oct 25;7:32 PMID: 17961241
  20. Practical implementation of an existing smoking detection pipeline and reduced support vector machine training corpus requirements.
    J Am Med Inform Assoc. 2014 Jan-Feb;21(1):27-30 PMID: 23921192
  21. Facilitating pharmacogenetic studies using electronic health records and natural-language processing: a case study of warfarin.
    J Am Med Inform Assoc. 2011 Jul-Aug;18(4):387-91 PMID: 21672908
  22. Automatic classification of mammography reports by BI-RADS breast tissue composition class.
    J Am Med Inform Assoc. 2012 Sep-Oct;19(5):913-6 PMID: 22291166
  23. Determining word sequence variation patterns in clinical documents using multiple sequence alignment.
    AMIA Annu Symp Proc. 2011;2011:934-43 PMID: 22195152
  24. Identifying QT prolongation from ECG impressions using a general-purpose Natural Language Processor.
    Int J Med Inform. 2009 Apr;78 Suppl 1:S34-42 PMID: 18938105
  25. Text classification for assisting moderators in online health communities.
    J Biomed Inform. 2013 Dec;46(6):998-1005 PMID: 24025513
  26. Extracting principal diagnosis, co-morbidity and smoking status for asthma research: evaluation of a natural language processing system.
    BMC Med Inform Decis Mak. 2006 Jul 26;6:30 PMID: 16872495
Article Info
Journal
Journal of the American Medical Informatics Association : JAMIA
Abbr.
J Am Med Inform Assoc
ISSN
1527-974X
Published
2014-00-00
Epub
2014-00-27
Pages
850-7
Language
English
Region
England
NLM ID
9430800
PMCID
PMC4147608
Subset
IM
Grants
Canadian Institutes of Health Research · HIR 08-374 · Canada
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]