An Experimental Comparison of Naive Bayesian and Keyword-Based Anti-Spam Filtering with Personal E-mail Messages

dc.creatorAndroutsopoulos, Ion
dc.creatorKoutsias, John
dc.creatorChandrinos, Konstantinos V.
dc.creatorSpyropoulos, Constantine D.
dc.date2000-08-22
dc.date.accessioned2026-07-07T03:16:28Z
dc.date.available2026-07-07T03:16:28Z
dc.descriptionThe growing problem of unsolicited bulk e-mail, also known as "spam", has generated a need for reliable anti-spam e-mail filters. Filters of this type have so far been based mostly on manually constructed keyword patterns. An alternative approach has recently been proposed, whereby a Naive Bayesian classifier is trained automatically to detect spam messages. We test this approach on a large collection of personal e-mail messages, which we make publicly available in "encrypted" form contributing towards standard benchmarks. We introduce appropriate cost-sensitive measures, investigating at the same time the effect of attribute-set size, training-corpus size, lemmatization, and stop lists, issues that have not been explored in previous experiments. Finally, the Naive Bayesian filter is compared, in terms of performance, to a filter that uses keyword patterns, and which is part of a widely used e-mail reader.
dc.identifierhttps://arxiv.org/abs/cs/0008019
dc.identifierhttp://arxiv.org/abs/cs/0008019
dc.identifierProceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval, N.J. Belkin, P. Ingwersen and M.-K. Leong (Eds.), Athens, Greece, July 24-28, 2000, pages 160-167
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/30367
dc.subjectComputation and Language
dc.subjectInformation Retrieval
dc.subjectMachine Learning
dc.subjectH.4.3; I.2.6; I.2.7; I.5.4; K.4.1
dc.titleAn Experimental Comparison of Naive Bayesian and Keyword-Based Anti-Spam Filtering with Personal E-mail Messages
dc.typetext

Files

Collections