A Simple Mechanism for Focused Web-harvesting
| dc.creator | Akbar, Z. | |
| dc.creator | Handoko, L. T. | |
| dc.date | 2008-09-03 | |
| dc.date.accessioned | 2026-07-07T10:00:37Z | |
| dc.date.available | 2026-07-07T10:00:37Z | |
| dc.description | The focused web-harvesting is deployed to realize an automated and comprehensive index databases as an alternative way for virtual topical data integration. The web-harvesting has been implemented and extended by not only specifying the targeted URLs, but also predefining human-edited harvesting parameters to improve the speed and accuracy. The harvesting parameter set comprises three main components. First, the depth-scale of being harvested final pages containing desired information counted from the first page at the targeted URLs. Secondly, the focus-point number to determine the exact box containing relevant information. Lastly, the combination of keywords to recognize encountered hyperlinks of relevant images or full-texts embedded in those final pages. All parameters are accessible and fully customizable for each target by the administrators of participating institutions over an integrated web interface. A real implementation to the Indonesian Scientific Index which covers all scientific information across Indonesia is also briefly introduced. | |
| dc.description | 6 pages, 4 figures, Proceeding of the International Conference on Advanced Computational Intelligence and Its Applications 2008 | |
| dc.identifier | https://arxiv.org/abs/0809.0723 | |
| dc.identifier | http://arxiv.org/abs/0809.0723 | |
| dc.identifier.uri | http://salesiana.dossiersoluciones.com/handle/123456789/168367 | |
| dc.subject | Information Retrieval | |
| dc.subject | Computers and Society | |
| dc.title | A Simple Mechanism for Focused Web-harvesting | |
| dc.type | text |