Skip to main content

Preliminary investigations into the word categorization system of BERT, 2023

 Item — Call Number: MU Thesis Chi
Identifier: b7931714

Scope and Contents

From the Collection:

The collection consists of theses written by students enrolled in the Monmouth University graduate Computer Science program. The holdings are primarily bound print documents that were submitted in partial fulfillment of requirements for the Master of Science degree.

Dates

  • Creation: 2023

Creator

Language of Materials

From the Collection:

Unless noted otherwise at the resource component level, the language of the collection materials is English.

Conditions Governing Access

All analog collection holdings are limited to library use only.

Collection holdings may not be borrowed through Interlibrary Loan.

Researchers seeking to photocopy collection materials must complete an Application to Photocopy Form.

Any photocopying of collection materials will be performed by the Monmouth University Library staff.

The Monmouth University Library reserves the right to limit or refuse duplication requests subject to the condition of collection materials and/or restrictions imposed by the collection creators or by the United States Copyright Act.

Permission to examine, or copy, collection materials does not imply permission to publish or quote. It is the responsibility of the researcher to obtain such permissions from both the copyright holder and Monmouth University.

Full Extent

1 Items (print book) : 70 pages ; 8.5 x 11.0 inches (28 cm).

Abstract

Bidirectional Encoder Representations from Transformers (BERT), introduced by Google, is a powerful natural language processing model as it is able to understand the meaning of words in a sentence in context. WordNet, developed at Princeton University, is a lexical database that shows semantic relationships between words. This thesis looks to investigate BERT’s word categorization system by looking at groups of example sentences given from related WordNet synsets. Because BERT allows a contextual word embedding to be extracted from each sentence, this was done with each of the words that share a meaning in each of the sentences. Agglomerative hierarchical clustering was then used on these extracted embeddings to see how the words with similar meanings in each of the sentences were related. The results from clustering these sentences were then compared to how WordNet considered the sentences to be related.

Keywords: BERT, WordNet, natural language processing, agglomerative clustering, transformers, word embeddings.

Partial Contents

Abstract -- Acknowledgements -- List of figures -- Table of contents -- 1. Introduction -- 2. Background -- 3. Implementation and results -- 4. Future research -- 5. Conclusion -- 6. References -- 7. Appendix.

Repository Details

Part of the Monmouth University Library Archives Repository

Contact:
Monmouth University Library
400 Cedar Avenue
West Long Branch New Jersey 07764 United States
732-923-4526