Cross-lingual keyword assignment

dc.creatorSteinberger, Ralf
dc.date2006-09-12
dc.date.accessioned2026-07-07T07:23:51Z
dc.date.available2026-07-07T07:23:51Z
dc.descriptionThis paper presents a language-independent approach to controlled vocabulary keyword assignment using the EUROVOC thesaurus. Due to the multilingual nature of EUROVOC, the keywords for a document written in one language can be displayed in all eleven official European Union languages. The mapping of documents written in different languages to the same multilingual thesaurus furthermore allows cross-language document comparison. The assignment of the controlled vocabulary thesaurus descriptors is achieved by applying a statistical method that uses a collection of manually indexed documents to identify, for each thesaurus descriptor, a large number of lemmas that are statistically associated to the descriptor. These associated words are then used during the assignment procedure to identify a ranked list of those EUROVOC terms that are most likely to be good keywords for a given document. The paper also describes the challenges of this task and discusses the achieved results of the fully functional prototype.
dc.descriptionPrecursor paper to cs.CL/0609059. The automatic classification system described here has now matured and is in daily use for document indexing in a European parliament. See http://langtech.jrc.it/Eurovoc.html for more details. 8 pages
dc.identifierhttps://arxiv.org/abs/cs/0609061
dc.identifierhttp://arxiv.org/abs/cs/0609061
dc.identifierProceedings of the XVII Conference of the Spanish Society for Natural Language Processing (SEPLN-2001). Procesamiento del Lenguaje Natural, Revista No. 27, pp. 273-280. Jaen, Spain, 12-14 September 2001. ISSN 1135-5948
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/116137
dc.subjectComputation and Language
dc.subjectInformation Retrieval
dc.subjectH.3.1; H.3.3; H.3.4; H.3.6
dc.titleCross-lingual keyword assignment
dc.typetext

Files

Collections