A Simple Mechanism for Focused Web-harvesting

dc.creatorAkbar, Z.
dc.creatorHandoko, L. T.
dc.date2008-09-03
dc.date.accessioned2026-07-07T10:00:37Z
dc.date.available2026-07-07T10:00:37Z
dc.descriptionThe focused web-harvesting is deployed to realize an automated and comprehensive index databases as an alternative way for virtual topical data integration. The web-harvesting has been implemented and extended by not only specifying the targeted URLs, but also predefining human-edited harvesting parameters to improve the speed and accuracy. The harvesting parameter set comprises three main components. First, the depth-scale of being harvested final pages containing desired information counted from the first page at the targeted URLs. Secondly, the focus-point number to determine the exact box containing relevant information. Lastly, the combination of keywords to recognize encountered hyperlinks of relevant images or full-texts embedded in those final pages. All parameters are accessible and fully customizable for each target by the administrators of participating institutions over an integrated web interface. A real implementation to the Indonesian Scientific Index which covers all scientific information across Indonesia is also briefly introduced.
dc.description6 pages, 4 figures, Proceeding of the International Conference on Advanced Computational Intelligence and Its Applications 2008
dc.identifierhttps://arxiv.org/abs/0809.0723
dc.identifierhttp://arxiv.org/abs/0809.0723
dc.identifier.urihttp://salesiana.dossiersoluciones.com/handle/123456789/168367
dc.subjectInformation Retrieval
dc.subjectComputers and Society
dc.titleA Simple Mechanism for Focused Web-harvesting
dc.typetext

Files

Collections