Mining the Web for Synonyms: PMI-IR versus LSA on TOEFL
| dc.creator | Turney, Peter D. | |
| dc.date | 2002-12-11 | |
| dc.date.accessioned | 2026-07-07T03:19:16Z | |
| dc.date.available | 2026-07-07T03:19:16Z | |
| dc.description | This paper presents a simple unsupervised learning algorithm for recognizing synonyms, based on statistical data acquired by querying a Web search engine. The algorithm, called PMI-IR, uses Pointwise Mutual Information (PMI) and Information Retrieval (IR) to measure the similarity of pairs of words. PMI-IR is empirically evaluated using 80 synonym test questions from the Test of English as a Foreign Language (TOEFL) and 50 synonym test questions from a collection of tests for students of English as a Second Language (ESL). On both tests, the algorithm obtains a score of 74%. PMI-IR is contrasted with Latent Semantic Analysis (LSA), which achieves a score of 64% on the same 80 TOEFL questions. The paper discusses potential applications of the new unsupervised learning algorithm and some implications of the results for LSA and LSI (Latent Semantic Indexing). | |
| dc.description | 12 pages | |
| dc.identifier | https://arxiv.org/abs/cs/0212033 | |
| dc.identifier | http://arxiv.org/abs/cs/0212033 | |
| dc.identifier | Proceedings of the Twelfth European Conference on Machine Learning, (2001), Freiburg, Germany, 491-502 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/31392 | |
| dc.subject | Machine Learning | |
| dc.subject | Computation and Language | |
| dc.subject | Information Retrieval | |
| dc.subject | I.2.6; I.2.7; H.3.1; H.3.3 | |
| dc.title | Mining the Web for Synonyms: PMI-IR versus LSA on TOEFL | |
| dc.type | text |