Scientific Data Management in the Coming Decade
| dc.creator | Gray, Jim | |
| dc.creator | Liu, David T. | |
| dc.creator | Nieto-Santisteban, Maria | |
| dc.creator | Szalay, Alexander S. | |
| dc.creator | DeWitt, David | |
| dc.creator | Heber, Gerd | |
| dc.date | 2005-02-02 | |
| dc.date.accessioned | 2026-07-07T03:22:27Z | |
| dc.date.available | 2026-07-07T03:22:27Z | |
| dc.description | This is a thought piece on data-intensive science requirements for databases and science centers. It argues that peta-scale datasets will be housed by science centers that provide substantial storage and processing for scientists who access the data via smart notebooks. Next-generation science instruments and simulations will generate these peta-scale datasets. The need to publish and share data and the need for generic analysis and visualization tools will finally create a convergence on common metadata standards. Database systems will be judged by their support of these metadata standards and by their ability to manage and access peta-scale datasets. The procedural stream-of-bytes-file-centric approach to data analysis is both too cumbersome and too serial for such large datasets. Non-procedural query and analysis of schematized self-describing data is both easier to use and allows much more parallelism. | |
| dc.identifier | https://arxiv.org/abs/cs/0502008 | |
| dc.identifier | http://arxiv.org/abs/cs/0502008 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/32599 | |
| dc.subject | Databases | |
| dc.subject | Computational Engineering, Finance, and Science | |
| dc.title | Scientific Data Management in the Coming Decade | |
| dc.type | text |