Text mining with enhanced named entity recognition, 2017
Scope and Contents
The collection consists of theses written by students enrolled in the Monmouth University graduate Computer Science program. The holdings are primarily bound print documents that were submitted in partial fulfillment of requirements for the Master of Science degree.
Dates
- Creation: 2017
Creator
- Mallojula, Prashanthi (Author, Person)
- Scherl, Richard (Thesis advisor, Person)
Conditions Governing Access
All analog collection holdings are limited to library use only.
Collection holdings may not be borrowed through Interlibrary Loan.
Researchers seeking to photocopy collection materials must complete an Application to Photocopy Form.
Any photocopying of collection materials will be performed by the Monmouth University Library staff.
The Monmouth University Library reserves the right to limit or refuse duplication requests subject to the condition of collection materials and/or restrictions imposed by the collection creators or by the United States Copyright Act.
Permission to examine, or copy, collection materials does not imply permission to publish or quote. It is the responsibility of the researcher to obtain such permissions from both the copyright holder and Monmouth University.
Full Extent
1 Items (print book) : 58 pages ; 8.5 x 11.0 inches (28 cm).
Language of Materials
English
Abstract
The goal of this research is to evaluate the usefulness of enhanced named entity recognition for text mining. Text mining is a subpart of data mining. It is the application of data mining techniques to texts in natural language such as English. The data for this project consists news [sic] articles and article titles extracted from the Web. Enhanced name-entity recognition is used to add additional information to the text. Named entities in the text are recognized using the Stanford NER trigger. This identifies phrases that refer to persons, locations, or organizations and tags them with labels indicating the category to which they belong. These strings are used in fetching information from selected DBpedia archive data files. DBpedia is a semantically organized database that is automatically created from the structured content of Wikipedia. Text classification techniques are applied on news article text and article titles with and without the data from enhanced named entity recognition. The classification techniques Naïve Bayes, Decision tree and Random forest are used for text classification in this research. The performance of each technique is evaluated to find the improvement in text classification with enhanced name entity recognition.
Keywords: Text mining, classification, named entity recognition, data mining, Stanford NER tagger, DBpedia, Wikipedia, Naïve Bayes, decision tree, random forest, text classification.
Partial Contents
Abstract -- Acknowledgements -- 1. Introduction -- 2. News extraction Python -- 3. Named-entity recognizer (NER) -- 4. Categorization with DBpedia -- 5. Data preprocessing -- 6. Text classification -- 7. Classifying articles and titles -- Summary -- Bibliography -- Appendix.
Repository Details
Part of the Monmouth University Library Archives Repository
Monmouth University Library
400 Cedar Avenue
West Long Branch New Jersey 07764 United States
732-923-4526