Conserving Fuel in Statistical Language Learning: Predicting Data Requirements
| dc.creator | Lauer, Mark | |
| dc.date | 1995-09-07 | |
| dc.date.accessioned | 2026-07-07T09:10:00Z | |
| dc.date.available | 2026-07-07T09:10:00Z | |
| dc.description | In this paper I address the practical concern of predicting how much training data is sufficient for a statistical language learning system. First, I briefly review earlier results and show how these can be combined to bound the expected accuracy of a mode-based learner as a function of the volume of training data. I then develop a more accurate estimate of the expected accuracy function under the assumption that inputs are uniformly distributed. Since this estimate is expensive to compute, I also give a close but cheaply computable approximation to it. Finally, I report on a series of simulations exploring the effects of inputs that are not uniformly distributed. Although these results are based on simplistic assumptions, they are a tentative step toward a useful theory of data requirements for SLL systems. | |
| dc.description | 8 pages | |
| dc.identifier | https://arxiv.org/abs/cmp-lg/9509002 | |
| dc.identifier | http://arxiv.org/abs/cmp-lg/9509002 | |
| dc.identifier | Eighth Australian Joint Conference on Artificial Intelligence, Canberra, 1995. | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/151235 | |
| dc.subject | Computation and Language | |
| dc.title | Conserving Fuel in Statistical Language Learning: Predicting Data Requirements | |
| dc.type | text |