Bootstrapping Lexical Choice via Multiple-Sequence Alignment
| dc.creator | Barzilay, Regina | |
| dc.creator | Lee, Lillian | |
| dc.date | 2002-05-25 | |
| dc.date.accessioned | 2026-07-07T03:18:27Z | |
| dc.date.available | 2026-07-07T03:18:27Z | |
| dc.description | An important component of any generation system is the mapping dictionary, a lexicon of elementary semantic expressions and corresponding natural language realizations. Typically, labor-intensive knowledge-based methods are used to construct the dictionary. We instead propose to acquire it automatically via a novel multiple-pass algorithm employing multiple-sequence alignment, a technique commonly used in bioinformatics. Crucially, our method leverages latent information contained in multi-parallel corpora -- datasets that supply several verbalizations of the corresponding semantics rather than just one. We used our techniques to generate natural language versions of computer-generated mathematical proofs, with good results on both a per-component and overall-output basis. For example, in evaluations involving a dozen human judges, our system produced output whose readability and faithfulness to the semantic input rivaled that of a traditional generation system. | |
| dc.description | 8 pages; to appear in the proceedings of EMNLP-2002 | |
| dc.identifier | https://arxiv.org/abs/cs/0205065 | |
| dc.identifier | http://arxiv.org/abs/cs/0205065 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/31116 | |
| dc.subject | Computation and Language | |
| dc.subject | 1.2.7 | |
| dc.title | Bootstrapping Lexical Choice via Multiple-Sequence Alignment | |
| dc.type | text |