Plagiarism Detection in arXiv

dc.creatorSorokina, Daria
dc.creatorGehrke, Johannes
dc.creatorWarner, Simeon
dc.creatorGinsparg, Paul
dc.date2007-02-01
dc.date.accessioned2026-07-07T07:44:12Z
dc.date.available2026-07-07T07:44:12Z
dc.descriptionWe describe a large-scale application of methods for finding plagiarism in research document collections. The methods are applied to a collection of 284,834 documents collected by arXiv.org over a 14 year period, covering a few different research disciplines. The methodology efficiently detects a variety of problematic author behaviors, and heuristics are developed to reduce the number of false positives. The methods are also efficient enough to implement as a real-time submission screen for a collection many times larger.
dc.descriptionSixth International Conference on Data Mining (ICDM'06), Dec 2006
dc.identifierhttps://arxiv.org/abs/cs/0702012
dc.identifierhttp://arxiv.org/abs/cs/0702012
dc.identifierdoi:10.1109/ICDM.2006.126
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/123113
dc.subjectDatabases
dc.subjectDigital Libraries
dc.subjectInformation Retrieval
dc.titlePlagiarism Detection in arXiv
dc.typetext

Files

Collections