Assessing agreement on classification tasks: the kappa statistic

dc.creatorCarletta, Jean
dc.date1996-02-27
dc.date.accessioned2026-07-07T09:10:07Z
dc.date.available2026-07-07T09:10:07Z
dc.descriptionCurrently, computational linguists and cognitive scientists working in the area of discourse and dialogue argue that their subjective judgments are reliable using several different statistics, none of which are easily interpretable or comparable to each other. Meanwhile, researchers in content analysis have already experienced the same difficulties and come up with a solution in the kappa statistic. We discuss what is wrong with reliability measures as they are currently used for discourse and dialogue work in computational linguistics and cognitive science, and argue that we would be better off as a field adopting techniques from content analysis.
dc.description9 pages
dc.identifierhttps://arxiv.org/abs/cmp-lg/9602004
dc.identifierhttp://arxiv.org/abs/cmp-lg/9602004
dc.identifierComputational Lingustics 22:2 (1996 forthcoming)
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/151264
dc.subjectComputation and Language
dc.titleAssessing agreement on classification tasks: the kappa statistic
dc.typetext

Files

Collections