Empirical Methods for Compound Splitting

dc.creatorKoehn, Philipp
dc.creatorKnight, Kevin
dc.date2003-02-22
dc.date.accessioned2026-07-07T03:19:28Z
dc.date.available2026-07-07T03:19:28Z
dc.descriptionCompounded words are a challenge for NLP applications such as machine translation (MT). We introduce methods to learn splitting rules from monolingual and parallel corpora. We evaluate them against a gold standard and measure their impact on performance of statistical MT systems. Results show accuracy of 99.1% and performance gains for MT of 0.039 BLEU on a German-English noun phrase translation task.
dc.description8 pages, 2 figures. Published at EACL 2003
dc.identifierhttps://arxiv.org/abs/cs/0302032
dc.identifierhttp://arxiv.org/abs/cs/0302032
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/31477
dc.subjectComputation and Language
dc.subjectI.2.7
dc.titleEmpirical Methods for Compound Splitting
dc.typetext

Files

Collections