Home LiteratureArticle Details
PMID: 27044929 Published · ppublish English Journal Article

PDF text classification to leverage information extraction from publication reports.

Journal of biomedical informatics ·Vol. 61 ·2016-00-00 ·Pages 141-8

Bui DD, Del Fiol G, Jonnalagadda S

Abstract

Data extraction from original study reports is a time-consuming, error-prone process in systematic review development. Information extraction (IE) systems have the potential to assist humans in the extraction task, however majority of IE systems were not designed to work on Portable Document Format (PDF) document, an important and common extraction source for systematic review. In a PDF document, narrative content is often mixed with publication metadata or semi-structured text, which add challenges to the underlining natural language processing algorithm. Our goal is to categorize PDF texts for strategic use by IE systems. We used an open-source tool to extract raw texts from a PDF document and developed a text classification algorithm that follows a multi-pass sieve framework to automatically classify PDF text snippets (for brevity, texts) into TITLE, ABSTRACT, BODYTEXT, SEMISTRUCTURE, and METADATA categories. To validate the algorithm, we developed a gold standard of PDF reports that were included in the development of previous systematic reviews by the Cochrane Collaboration. In a two-step procedure, we evaluated (1) classification performance, and compared it with machine learning classifier, and (2) the effects of the algorithm on an IE system that extracts clinical outcome mentions. The multi-pass sieve algorithm achieved an accuracy of 92.6%, which was 9.7% (p<0.001) higher than the best performing machine learning classifier that used a logistic regression algorithm. F-measure improvements were observed in the classification of TITLE (+15.6%), ABSTRACT (+54.2%), BODYTEXT (+3.7%), SEMISTRUCTURE (+34%), and MEDADATA (+14.2%). In addition, use of the algorithm to filter semi-structured texts and publication metadata improved performance of the outcome extraction system (F-measure +4.1%, p=0.002). It also reduced of number of sentences to be processed by 44.9% (p<0.001), which corresponds to a processing time reduction of 50% (p=0.005). The rule-based multi-pass sieve framework can be used effectively in categorizing texts extracted from PDF documents. Text classification is an important prerequisite step to leverage information extraction from PDF documents.

Keywords
Document analysis Machine learning Natural language processing Text classification
MeSH Terms
Algorithms Humans Information Storage and Retrieval Machine Learning Narration Natural Language Processing Publications Review Literature as Topic
Authors & Affiliations
3 authors, click to expand affiliations / ORCID
Bui Duy Duc An
Department of Biomedical Informatics, University of Utah, Salt Lake City, UT, USA; Department of Preventive Medicine-Health and Biomedical Informatics, Northwestern University, Chicago, IL, USA. Electronic address: [email protected].
Del Fiol Guilherme
Department of Biomedical Informatics, University of Utah, Salt Lake City, UT, USA.
Jonnalagadda Siddhartha
Department of Preventive Medicine-Health and Biomedical Informatics, Northwestern University, Chicago, IL, USA.
References (23)
23 references, click to expand
  1. Learning regular expressions for clinical text classification.
    J Am Med Inform Assoc. 2014 Sep-Oct;21(5):850-7 PMID: 24578357
  2. Coreference analysis in clinical notes: a multi-pass sieve with alternate anaphora resolution modules.
    J Am Med Inform Assoc. 2012 Sep-Oct;19(5):867-74 PMID: 22707745
  3. How quickly do systematic reviews go out of date? A survival analysis.
    Ann Intern Med. 2007 Aug 21;147(4):224-33 PMID: 17638714
  4. Automatic extracting of patient-related attributes: disease, age, gender and race.
    Stud Health Technol Inform. 2012;180:589-93 PMID: 22874259
  5. Automated extraction of reported statistical analyses: towards a logical representation of clinical trial literature.
    AMIA Annu Symp Proc. 2012;2012:350-9 PMID: 23304305
  6. Automatically finding relevant citations for clinical guideline development.
    J Biomed Inform. 2015 Oct;57:436-45 PMID: 26363352
  7. PICO element detection in medical text without metadata: are first sentences enough?
    J Biomed Inform. 2013 Oct;46(5):940-6 PMID: 23899909
  8. Combining classifiers for robust PICO element detection.
    BMC Med Inform Decis Mak. 2010 May 15;10:29 PMID: 20470429
  9. Biomedical literature classification using encyclopedic knowledge: a Wikipedia-based bag-of-concepts approach.
    PeerJ. 2015 Sep 29;3:e1279 PMID: 26468436
  10. Evidence based medicine: what it is and what it isn't.
    BMJ. 1996 Jan 13;312(7023):71-2 PMID: 8555924
  11. Support Vector Feature Selection for Early Detection of Anastomosis Leakage From Bag-of-Words in Electronic Health Records.
    IEEE J Biomed Health Inform. 2016 Sep;20(5):1404-15 PMID: 25312965
  12. Efficient extraction of protein-protein interactions from full-text articles.
    IEEE/ACM Trans Comput Biol Bioinform. 2010 Jul-Sep;7(3):481-94 PMID: 20498514
  13. High prevalence but low impact of data extraction and reporting errors were found in Cochrane systematic reviews.
    J Clin Epidemiol. 2005 Jul;58(7):741-2 PMID: 15939227
  14. A hybrid approach to sentiment sentence classification in suicide notes.
    Biomed Inform Insights. 2012;5(Suppl. 1):43-50 PMID: 22879759
  15. Hidden Markov models.
    Curr Opin Struct Biol. 1996 Jun;6(3):361-5 PMID: 8804822
  16. Classification of diffuse lung disease patterns on high-resolution computed tomography by a bag of words approach.
    Med Image Comput Comput Assist Interv. 2011;14(Pt 3):183-90 PMID: 22003698
  17. BioRAT: extracting biological information from full-length papers.
    Bioinformatics. 2004 Nov 22;20(17):3206-13 PMID: 15231534
  18. A knowledge engineering approach to recognizing and extracting sequences of nucleic acids from scientific literature.
    Conf Proc IEEE Eng Med Biol Soc. 2010;2010:1081-4 PMID: 21096556
  19. Automated information extraction of key trial design elements from clinical trial publications.
    AMIA Annu Symp Proc. 2008 Nov 06;:141-5 PMID: 18999067
  20. A multi-classifier based guideline sentence classification system.
    Healthc Inform Res. 2011 Dec;17(4):224-31 PMID: 22259724
  21. ExaCT: automatic extraction of clinical trial characteristics from journal publications.
    BMC Med Inform Decis Mak. 2010 Sep 28;10:56 PMID: 20920176
  22. Detection of protein catalytic sites in the biomedical literature.
    Pac Symp Biocomput. 2013;:433-44 PMID: 23424147
  23. The Global Evidence Mapping Initiative: scoping research in broad topic areas.
    BMC Med Res Methodol. 2011 Jun 17;11:92 PMID: 21682870
Article Info
Journal
Journal of biomedical informatics
Abbr.
J Biomed Inform
ISSN
1532-0480
Published
2016-00-00
Epub
2016-00-01
Pages
141-8
Language
English
Region
United States
NLM ID
100970413
PMCID
PMC4893911
Subset
IM
Grants
NLM NIH HHS · R00 LM011389 · United States
NLM NIH HHS · R01 LM011416 · United States
Analysis Services
Analysis Services

Contact

No. 2 Wenbo Road, Zhangqiu District, Jinan, Shandong

Qilu Normal University · Genelibs Bioinformatics Lab

750 Shunhua Rd, Jinan

2F, Bldg F, University Science Park

Tel: 0531-88819269

WeChat Official Account

Follow our WeChat subscription account for real-time updates and the latest in medical and biological research.


Business Email

E-mail: [email protected]