Tagger Evaluation Given Hierarchical Tag Sets

dc.creatorMelamed, I. Dan
dc.creatorResnik, Philip
dc.date2000-08-10
dc.date.accessioned2026-07-07T03:16:26Z
dc.date.available2026-07-07T03:16:26Z
dc.descriptionWe present methods for evaluating human and automatic taggers that extend current practice in three ways. First, we show how to evaluate taggers that assign multiple tags to each test instance, even if they do not assign probabilities. Second, we show how to accommodate a common property of manually constructed ``gold standards'' that are typically used for objective evaluation, namely that there is often more than one correct answer. Third, we show how to measure performance when the set of possible tags is tree-structured in an IS-A hierarchy. To illustrate how our methods can be used to measure inter-annotator agreement, we show how to compute the kappa coefficient over hierarchical tag sets.
dc.descriptionpreprint is 7 pages, laid out differently than printed version
dc.identifierhttps://arxiv.org/abs/cs/0008007
dc.identifierhttp://arxiv.org/abs/cs/0008007
dc.identifierComputers and the Humanities 34(1-2). Special issue on SENSEVAL. pp. 79-84
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/30355
dc.subjectComputation and Language
dc.subjectG.3; I.2.7; J.5
dc.titleTagger Evaluation Given Hierarchical Tag Sets
dc.typetext

Files

Collections