Error-Driven Pruning of Treebank Grammars for Base Noun Phrase Identification

dc.creatorCardie, Claire
dc.creatorPierce, David
dc.date1998-08-26
dc.date.accessioned2026-07-07T02:36:22Z
dc.date.available2026-07-07T02:36:22Z
dc.descriptionFinding simple, non-recursive, base noun phrases is an important subtask for many natural language processing applications. While previous empirical methods for base NP identification have been rather complex, this paper instead proposes a very simple algorithm that is tailored to the relative simplicity of the task. In particular, we present a corpus-based approach for finding base NPs by matching part-of-speech tag sequences. The training phase of the algorithm is based on two successful techniques: first the base NP grammar is read from a ``treebank'' corpus; then the grammar is improved by selecting rules with high ``benefit'' scores. Using this simple algorithm with a naive heuristic for matching rules, we achieve surprising accuracy in an evaluation on the Penn Treebank Wall Street Journal.
dc.description7 pages; 2 eps figures; uses epsf, colacl
dc.identifierhttps://arxiv.org/abs/cmp-lg/9808015
dc.identifierhttp://arxiv.org/abs/cmp-lg/9808015
dc.identifierProceedings of COLING-ACL'98, pages 218-224.
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/15890
dc.subjectComputation and Language
dc.titleError-Driven Pruning of Treebank Grammars for Base Noun Phrase Identification
dc.typetext

Files

Collections