M.Phil / MS
Permanent URI for this collectionhttps://repository.cuilahore.edu.pk/handle/123456789/36
This collection archives the complete set of theses produced by students of the COMSATS University Islamabad, Lahore Campus.
Browse
5 results
Search Results
Item Fake News Detection and Corpus Development for Urdu News(Library Information Services, COMSATS University Islamabad, Lahore Campus, 2021) Muhammad Ismail Suri; SP17-RCS-009; LHR TP 7287; Dr. Muhammad Waqas AnwarWith the growing technological advancements in the present era, the curiosity to remain updated about the daily happenings in the world has opened up various opportunities for news dissemination. Access to information with ease has made the propagation of news much efficient and cost effective. However, online news is a double-edged sword that creates a room for spreading fake news which intends to mislead and exaggerate the reality. This issue is a rising problem leading towards several real world sociopolitical and cultural impacts. To curb the menace of fake news different organizations are working to handle it, but majority of the work is being done in English, unfortunately the language of Urdu in this dimension remains untouched. This research is focused on detecting fake news in Urdu language. The demand is to have a detecting mechanism that can predict between fake and real news. In South Asian languages, unavailability of datasets is the major hindrance of evaluating fake news. Therefore, my contribution is the induction of a dataset influenced by English fake news dataset for Urdu news; Corpus of Urdu News for Fake News Detection (COUNFND) which has been derived from digital journalism. It contains 480 documents with 120 claims processed by means of conducting traditional machine learning algorithms and text preprocessing techniques. The corpus holds great importance for the evaluation of fake news detection system specifically for Urdu Language. This thesis discusses different techniques that have been applied on the Urdu language datasets and furthermore achieving classification accuracy of about 73 percent effectively. Later on, a new dataset of fake news in Urdu Language has been introduced by Maaz Amjad (Bend the Truth), which has articles divided into fake and real classification, I have applied different supervised learning models and have achieved an accuracy of 0.89 on random forest using four grams techniques which is greater than the baseline accuracy provided with the dataset.Item Group Spam identification in Online Product Reviews(Library Information Services, COMSATS University Islamabad, Lahore Campus, 2020) GHULAM MEHMOOD; SP17-RCS-017; LHR TP 6400; Dr. Hamid Turab MirzaIn this era of e-commerce, user’s opinion about a product on the online review system is of great importance, as it provides guidance for people to decide. Being very important, people used to write fake reviews about the products, called opinion spamming. Detecting opinion spams in online review platforms is a challenging task drawing attention from research communities. It is a persistent campaign between the spammers and platforms. Grouped opinion spamming is the main type of opinion spamming in the online review system these days. For that purpose, usually people make multiple accounts to write fake reviews or they pay to crowdsourcing platforms to write fake reviews for promotion of their product or demotion of their competitor’s product. Group spam reviews are more damaging for online review systems as reviews from many peoples about a product either it is positive, or negative can easily deceive peoples as compared to a single spam review. These spam review groups should be detected so that reviews about that product reflect genuine user opinion. Many researchers try to resolve this problem using behavioral and linguistic features of the users and reviews. Many machine learning models are being adapted to solve this problem but could not resolve this problem completely. The purpose of this research work is to design a framework that detect group spammer who targets online review systems. This framework has used linguistic, behavioral, and structural features to dig out all such spammer groups who write fake reviews. First, constructed a reviewer-product network between reviewers and products. Then written an algorithm to find out collusiveness score between reviewers who have commonly reviewed a product and constructed a network between reviewers. Then all those edges (reviewer pairs) whose score was below a threshold valued was removed from the network. After that extracted high quality candidate spammer groups by using a novel algorithm. Five best group spam indicators are used to calculate spamicity score of candidate spammer groups. The candidate groups whose spamicity score was greater than a threshold values (says 0.6) was considered as spammer groups. For the evaluation of proposed algorithm, experiments performed on 3 labeled datasets from Yelp. The system extracted spammer groups from online reviews with precision of 0.89 @ top 50 spammer groups.Item Deception Detection using Facial Action Coding System in Videos(Library Information Services, COMSATS University Islamabad, Lahore Campus, 2020) Hammad-ud-din Ahmed; SP17-RCS-015; LHR TP 5974; Dr. Usama Ijaz BajwaHumans lie to each other on a regular basis in one form or another. To lie to someone is to share altered facts. Facts are important for decision making in every situation. Without unaltered facts, potentially harmful decisions can be made which can alter someone’s life in ways that may not want. This is why it is important to catch misinformation before any harm can be caused by them. Deception detection in videos has gained traction in recent times for its various real-life applications. When a person lies, they show facial expressions, also known as microexpressions so fast that they can be ignored by the untrained eye. Microexpressions can be used as a basis for creating a deception detection system. Facial Action Coding System is utilized as a way to encode and extract data from facial muscle movement during truthful and deceptive confessions. This data is used to train a deep learning model that utilizes long short-term memory (LSTM) to train and create the system. The real-life trial dataset is used to train and test the system provided one of the best facial only approaches to deception detection. Cross-data validation is also tested using the Real-life trial dataset, the Silesian Deception Database, and the Bag-of-lies Deception Database. Cross-data validation is something that has not yet been attempted by anyone else for a deception detection system which provides unique insight for the field. The results show that adding different datasets to train a neural network for the sake of creating a deception detection system worsens the accuracy of the system but the reasoning behind the poorer results actually encourages attempts at creating better datasetsItem Interest Flooding Attack Mitigation in Vehicular Name Data Networks (VNDN)(Library Information Services, COMSATS University Islamabad, Lahore Campus, 2020) Muhammad Abdullah; SP17-RCS-004; LHR TP 5973; Professor Dr. Syed Asad HussainIn connected vehicular ecosystem, multimedia data sharing is rapidly increasing and it has emerged as a crucial aspect of Vehicular Ad-hoc Network (VANET). Traditional TCP/IP network model is found inappropriate to cope with high mobility, network volatility and intermittent connectivity between vehicular nodes. During the last few years, Named Data Network (NDN) has been considered as a promising candidate for efficient and seamless connectivity between vehicular nodes. Vehicular Named Data Network (VNDN) has been used to handle intermittent connectivity between vehicles, mobility, seamless connectivity and scalability problems in host-centric VANET environment. In VNDN, communication takes place on the basis of name instead of host IP address mechanism. Moreover, communication between vehicles follows pull based approach, i.e., data packets are generated on demand to prevent network from overloading. However, there are certain security challenges in a VNDN environment. These challenges are Denial-of-Service (DoS) Attacks, Cache related attacks, timing attacks, Interest Flooding Attack (IFA) etc. This research work focuses on Interest Flooding Attack (IFA) in a VNDN scenario. In IFA, attacker(s) generates malicious Interest packets with non-existent prefix to exhaust network resources and cause Interest packet flooding across the network. This thesis has proposed a novel attack mitigation scheme to counter this problem. The proposed Priority based per-flow Interest rate monitoring scheme determines suspicious flow of malicious incoming Interest packets in attacked vehicles. Moreover, proposed attack mitigation method comprises three different phases, i.e., attack detection, recovery and prevention, to prevent VNDN scenario from shutting down. Detection phase identifies suspicious flow of incoming Interest packets while attack recovery phase back tracks attacker vehicle(s). The proposed research scheme has introduced attack prevention phase to limit scalability of IFA in a VNDN scenario. A priority flag is assigned to incoming flow of Interest packets that detects IFA. The priority of incoming Interest packet flow is calculated on the basis of various parameters, i.e., Average Interest transmission rate, Interest Satisfaction Rate, Interest packet retransmission count, Interest Satisfaction Ratio (ISR), number of hop x traversed, cache hit/miss ratio and Interest packet transmission delay. The aforementioned metrics are considered to avoid false attack detection and rejection of legitimate Interest packets. Repeated Simulations show that proposed IFA mitigation scheme outperforms prior per face Interest rate limiting solutions in terms of improved Interest Satisfaction rate, cache hit ratio and ISR during attack window. Besides this, proposed research model also maintains low rate of Pending Interest Table (PIT) utilization, packet collisions, Interest packets retransmission count, end-to-end delay and ratio of Unsatisfied Interest packets. Furthermore, scalability of proposed research strategy is also evaluated by changing density of attackers at real-time. Moreover in proposed attack mitigation model, rate of incoming legitimate Interest packets increases by reducing drop rate of valid Interest packets.Item Toxic Comment Classification of Roman Urdu Text(Library Information Services, COMSATS University Islamabad, Lahore Campus, 2019) Waheed Abbas; FA17-RCS-013; LHR TP 5624; Dr. Rao Muhammad Adeel NawabWe are living in an era of technology where most of us have easy access to the inter net. Due to internet bloom, the use of social media and discussion forum, in particular for communication, increase tremendously in recent years. But this advancement also opens doors for trolls who poison these social media and forums by their abusive behav ior toward others. Detection of toxic language online is becoming an important issue in recent years. A toxic comment is defined as rude, unreasonable, or disrespectful com ments which likely to make you leave a conversation. Current methods for handling toxicity online are often heavily dependent on manual moderation, and not scalable enough to handle growing users day by day. In this thesis, we investigated the problem of toxic language detection for Roman Urdu (transliteration of Urdu in English) text because Roman Urdu is widely used to communicate online in Pakistani and Indian community. Various work has been done to detect toxicity in the English language, but due to unavailability of large Roman Urdu corpora, there is no work done so far in Roman Urdu. The aim of this thesis work is two folds: first is to develop large scale corpus for detection of toxic language in Roman Urdu, second is to apply different proven classical machine learning and state-of-the-art deep learning methods on the developed corpus. In the first step, we collected data from the YouTube comment section to develop Roman Urdu corpus. The final corpus consists of 10,863 unique comments where we have 5,887 toxic and 4,976 non-toxic comments. We performed manual annotation of corpus and got it reviewed from another annotator. We applied proven classical ma chine learning algorithms like Logistic Regression, K-Nearest Neighbors, Naive Bayes, Random Forest, Support Vector Machine and Multi-layer Perceptron along with state of-the-art deep learning methods like Long Short-Term Memory, Bidirectional Long Short-Term Memory, and Convolutional Neural Networks. Best results obtained from Random Forest with character 4-gram (Accuracy = 0.964, Precision = 0.980, Recall = 0.953, F1 = 0.966).