A Resource Efficient Method for Indexing Hidden Web Using Rank Based Web Crawling Techniques

dc.contributor.authorAbdul Mannan
dc.contributor.authorFA18-RCS-026
dc.contributor.authorLHR TP 7282
dc.contributor.authorDr. Muhammad Hasanain Chaudary
dc.date.accessioned2026-02-13T11:02:35Z
dc.date.issued2021
dc.description.abstractToday we are seeing a shift in understanding and behavior of individuals toward anonymity and privacy. As a result, not only usage of virtual private network is increasing, but also more and more people are converging towards hidden web. Hidden web is a server less chain-based architecture, which can only be accessed using specific proxies and gateways and provides anonymity-using chain of interconnected nodes over public IP so that even if node is compromised anonymity of user is maintained. This form of security has its own drawbacks as overall network speed depends on node with lowest connectivity. Subsequently, crawling becomes costly since speed is proportional to the latency of generated Tor circuit. Another challenge in crawling Hidden web is the volatile nature of hosted services provided on it. Due to its anonymous nature, illegal services are prevalent. Consequently, this network is highly monitored becoming hot bed of banned websites some of which become live on new nodes while other stay down. This causes loss of time and resources crawler used to mine those dead URL. We want to propose a crawler that can mine data in low latency Tor network, auto tuning its configuration according to the state of the network and giving rank to websites in line with their content. This rank would be used to calculate crawling depth of a specific service at a given time, further improving as service stays alive
dc.identifier.urihttps://repository.cuilahore.edu.pk/handle/123456789/1564
dc.language.isoen
dc.publisherLibrary Information Services, COMSATS University Islamabad, Lahore Campus
dc.relation.ispartofseriesLHR TP 7282
dc.subjectDr. Muhammad Hasanain Chaudary
dc.subjectFA18
dc.subjectComputer Science
dc.subjectDepartment of Computer Science
dc.subjectWeb Crawling Techniques
dc.titleA Resource Efficient Method for Indexing Hidden Web Using Rank Based Web Crawling Techniques
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
7282.docx
Size:
1.39 MB
Format:
Microsoft Word XML

License bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
319 B
Format:
Item-specific license agreed to upon submission
Description:

Collections