Automatic Identification of Document Translations in Large Multilingual Document Collections
| dc.creator | Pouliquen, Bruno | |
| dc.creator | Steinberger, Ralf | |
| dc.creator | Ignat, Camelia | |
| dc.date | 2006-09-12 | |
| dc.date.accessioned | 2026-07-07T07:23:51Z | |
| dc.date.available | 2026-07-07T07:23:51Z | |
| dc.description | Texts and their translations are a rich linguistic resource that can be used to train and test statistics-based Machine Translation systems and many other applications. In this paper, we present a working system that can identify translations and other very similar documents among a large number of candidates, by representing the document contents with a vector of thesaurus terms from a multilingual thesaurus, and by then measuring the semantic similarity between the vectors. Tests on different text types have shown that the system can detect translations with over 96% precision in a large search space of 820 documents or more. The system was tuned to ignore language-specific similarities and to give similar documents in a second language the same similarity score as equivalent documents in the same language. The application can also be used to detect cross-lingual document plagiarism. | |
| dc.description | This technology is used daily to link related news items across languages in the multilingual news analysis system NewsExplorer, which is freely accessible at http://press.jrc.it/NewsExplorer . 8 pages | |
| dc.identifier | https://arxiv.org/abs/cs/0609060 | |
| dc.identifier | http://arxiv.org/abs/cs/0609060 | |
| dc.identifier | Proceedings of the International Conference 'Recent Advances in Natural Language Processing' (RANLP'2003), pp. 401-408. Borovets, Bulgaria, 10 - 12 September 2003 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/116136 | |
| dc.subject | Computation and Language | |
| dc.subject | Information Retrieval | |
| dc.subject | H.3.1; H.3.3; H.3.4; H.3.6 | |
| dc.title | Automatic Identification of Document Translations in Large Multilingual Document Collections | |
| dc.type | text |