Improving the latent dirichlet allocation document model with WordNet
Document Type
Conference Proceeding
Publication Date
4-8-2010
Abstract
In the e-intelligence/counter-intelligence domain, actionable information must be extracted, filtered, and correlated from massive amounts of disparate often free text data. The usefulness of the information depends on how we accomplish these steps and present the most relevant information to the analyst. One method for extracting information from free text is Latent Dirichlet Allocation (LDA), a document categorization technique to classify documents into cohesive topics. Although LDA accounts for some implicit relationships such as synonymy (same meaning) it often ignores other semantic relationships such as polysemy (different meanings), hyponym (subordinate), and meronym (part of). To compensate for this deficiency, we incorporate explicit word ontologies, such as WordNet, into the LDA algorithm to account for various semantic relationships. Experiments over well-known document collections, 20 Newsgroups, NIPS, and OHSUMED, demonstrate that incorporating such background knowledge improves perplexity measure over LDA alone.
Source Publication
5th International Conference on Information Warfare and Security
Recommended Citation
Isaly, L., Trias, E. D., & Peterson, G. L. (2010). Improving the latent dirichlet allocation document model with wordnet. 5th International Conference on Information Warfare and Security, 163–170.
Comments
Current AFIT students, faculty and staff may access the full conference paper through ProQuest, by clicking here.