A Resource Efficient Method for Indexing Hidden Web Using Rank Based Web Crawling Techniques
| dc.contributor.author | Abdul Mannan | |
| dc.contributor.author | FA18-RCS-026 | |
| dc.contributor.author | LHR TP 7282 | |
| dc.contributor.author | Dr. Muhammad Hasanain Chaudary | |
| dc.date.accessioned | 2026-02-13T11:02:35Z | |
| dc.date.issued | 2021 | |
| dc.description.abstract | Today we are seeing a shift in understanding and behavior of individuals toward anonymity and privacy. As a result, not only usage of virtual private network is increasing, but also more and more people are converging towards hidden web. Hidden web is a server less chain-based architecture, which can only be accessed using specific proxies and gateways and provides anonymity-using chain of interconnected nodes over public IP so that even if node is compromised anonymity of user is maintained. This form of security has its own drawbacks as overall network speed depends on node with lowest connectivity. Subsequently, crawling becomes costly since speed is proportional to the latency of generated Tor circuit. Another challenge in crawling Hidden web is the volatile nature of hosted services provided on it. Due to its anonymous nature, illegal services are prevalent. Consequently, this network is highly monitored becoming hot bed of banned websites some of which become live on new nodes while other stay down. This causes loss of time and resources crawler used to mine those dead URL. We want to propose a crawler that can mine data in low latency Tor network, auto tuning its configuration according to the state of the network and giving rank to websites in line with their content. This rank would be used to calculate crawling depth of a specific service at a given time, further improving as service stays alive | |
| dc.identifier.uri | https://repository.cuilahore.edu.pk/handle/123456789/1564 | |
| dc.language.iso | en | |
| dc.publisher | Library Information Services, COMSATS University Islamabad, Lahore Campus | |
| dc.relation.ispartofseries | LHR TP 7282 | |
| dc.subject | Dr. Muhammad Hasanain Chaudary | |
| dc.subject | FA18 | |
| dc.subject | Computer Science | |
| dc.subject | Department of Computer Science | |
| dc.subject | Web Crawling Techniques | |
| dc.title | A Resource Efficient Method for Indexing Hidden Web Using Rank Based Web Crawling Techniques | |
| dc.type | Thesis |