Unsupervised Learning of Semantic Orientation from a Hundred-Billion-Word Corpus
| dc.creator | Turney, Peter D. | |
| dc.creator | Littman, Michael L. | |
| dc.date | 2002-12-08 | |
| dc.date.accessioned | 2026-07-07T03:19:13Z | |
| dc.date.available | 2026-07-07T03:19:13Z | |
| dc.description | The evaluative character of a word is called its semantic orientation. A positive semantic orientation implies desirability (e.g., "honest", "intrepid") and a negative semantic orientation implies undesirability (e.g., "disturbing", "superfluous"). This paper introduces a simple algorithm for unsupervised learning of semantic orientation from extremely large corpora. The method involves issuing queries to a Web search engine and using pointwise mutual information to analyse the results. The algorithm is empirically evaluated using a training corpus of approximately one hundred billion words -- the subset of the Web that is indexed by the chosen search engine. Tested with 3,596 words (1,614 positive and 1,982 negative), the algorithm attains an accuracy of 80%. The 3,596 test words include adjectives, adverbs, nouns, and verbs. The accuracy is comparable with the results achieved by Hatzivassiloglou and McKeown (1997), using a complex four-stage supervised learning algorithm that is restricted to determining the semantic orientation of adjectives. | |
| dc.description | 11 pages, issued 2002 | |
| dc.identifier | https://arxiv.org/abs/cs/0212012 | |
| dc.identifier | http://arxiv.org/abs/cs/0212012 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/31374 | |
| dc.subject | Machine Learning | |
| dc.subject | Information Retrieval | |
| dc.subject | H.3.1; H.3.3; I.2.6; I.2.7 | |
| dc.title | Unsupervised Learning of Semantic Orientation from a Hundred-Billion-Word Corpus | |
| dc.type | text |