Annotation graphs as a framework for multidimensional linguistic data analysis
| dc.creator | Bird, Steven | |
| dc.creator | Liberman, Mark | |
| dc.date | 1999-07-05 | |
| dc.date.accessioned | 2026-07-07T03:24:12Z | |
| dc.date.available | 2026-07-07T03:24:12Z | |
| dc.description | In recent work we have presented a formal framework for linguistic annotation based on labeled acyclic digraphs. These `annotation graphs' offer a simple yet powerful method for representing complex annotation structures incorporating hierarchy and overlap. Here, we motivate and illustrate our approach using discourse-level annotations of text and speech data drawn from the CALLHOME, COCONUT, MUC-7, DAMSL and TRAINS annotation schemes. With the help of domain specialists, we have constructed a hybrid multi-level annotation for a fragment of the Boston University Radio Speech Corpus which includes the following levels: segment, word, breath, ToBI, Tilt, Treebank, coreference and named entity. We show how annotation graphs can represent hybrid multi-level structures which derive from a diverse set of file formats. We also show how the approach facilitates substantive comparison of multiple annotations of a single signal based on different theoretical models. The discussion shows how annotation graphs open the door to wide-ranging integration of tools, formats and corpora. | |
| dc.description | 10 pages, 10 figures, Towards Standards and Tools for Discourse Tagging, Proceedings of the Workshop. pp. 1-10. Association for Computational Linguistics | |
| dc.identifier | https://arxiv.org/abs/cs/9907003 | |
| dc.identifier | http://arxiv.org/abs/cs/9907003 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/33233 | |
| dc.subject | Computation and Language | |
| dc.subject | A.1; E.2; H.2.1; H.3.3; H.3.7; I.2.7 | |
| dc.title | Annotation graphs as a framework for multidimensional linguistic data analysis | |
| dc.type | text |