Building a common pipeline for rule-based document classification

Patterson, O.V.; Ginter, T.; DuVall, S.L.

Studies in Health Technology and Informatics 192: 1211

2013


ISSN/ISBN: 1879-8365
PMID: 23920985
Document Number: 666019
Instance-based classification of clinical text is a widely used natural language processing task employed as a step for patient classification, document retrieval, or information extraction. Rule-based approaches rely on concept identification and context analysis in order to determine the appropriate class. We propose a five-step process that enables even small research teams to develop simple but powerful rule-based NLP systems by taking advantage of a common UIMA AS based pipeline for classification. Our proposed methodology coupled with the general-purpose solution provides researchers with access to the data locked in clinical text in cases of limited human resources and compact timelines.

Document emailed within 1 workday
Secure & encrypted payments