M.Phil / MS
Permanent URI for this collectionhttps://repository.cuilahore.edu.pk/handle/123456789/36
This collection archives the complete set of theses produced by students of the COMSATS University Islamabad, Lahore Campus.
Browse
2 results
Search Results
Item Fake News Detection for Urdu language using Machine Learning Techniques(Library Information Services, COMSATS University Islamabad, Lahore Campus, 2021) Ahmad Raza Khan; FA18-RCS-005; LHR TP 7587; Dr. Hasan JamalThe main target of this research is to detect fake news in Urdu language. Although many studies have already been conducted on fake news detection for different languages but not much work has been done for Urdu language and there is a research gap for fake news detection in Urdu language. To compare, evaluate and develop a fake news detection system for Urdu language, standardized and high-quality benchmark resources are needed but unfortunately corpora for Urdu language are not available, even though Urdu is a morphologically rich language and has more than 300 million speakers all over the world. To compensate for this gap, we develop a novel high quality, and a large-sized benchmark corpus for Urdu fake news detection. Only one benchmark dataset is previously available for fake news detection tasks in the Urdu language that contains a total of 900 news articles in which 500 are real news and 400 are fake news articles. Our novel corpus contains 1800 news articles out of which 1000 are real news articles and 800 are fake news articles. Our dataset contains five domain data (Business, Health, Showbiz, Sports, and Technology). These news articles are tagged manually into real and fake news tags. These news articles are collected from different authentic news sources. After corpus creation, feature extraction is performed using different combinations of word n-gram, character n-gram, and function n-gram. The range of this n-gram varies from 1 to 6 for our experiments. We use different weighting schemes such as Term Frequency-Inverse Document Frequency (TF-IDF), Log entropy, Binary weighting schemes, etc. After extracting features from the text, state-of-the-art machine learning techniques are applied to the corpus to evaluate the fake news detection system. In machine learning techniques, Multinomial Naïve Bayes (MNB), Linear Support Vector Classifier, Logistic regression, Random Forest, Decision Tree, Bernoulli Naïve Bayes (BNB) and Ada-Boost are used. Each classifier is trained and tested separately. Standard evaluation measures are used to check the working of the fake news detection system i.e., accuracy, precision, recall, F1-score. We compare our results with previous research results and Linear SVC and Logistic Regression outperforms other classifiers.Item Fake News Detection and Corpus Development for Urdu News(Library Information Services, COMSATS University Islamabad, Lahore Campus, 2021) Muhammad Ismail Suri; SP17-RCS-009; LHR TP 7287; Dr. Muhammad Waqas AnwarWith the growing technological advancements in the present era, the curiosity to remain updated about the daily happenings in the world has opened up various opportunities for news dissemination. Access to information with ease has made the propagation of news much efficient and cost effective. However, online news is a double-edged sword that creates a room for spreading fake news which intends to mislead and exaggerate the reality. This issue is a rising problem leading towards several real world sociopolitical and cultural impacts. To curb the menace of fake news different organizations are working to handle it, but majority of the work is being done in English, unfortunately the language of Urdu in this dimension remains untouched. This research is focused on detecting fake news in Urdu language. The demand is to have a detecting mechanism that can predict between fake and real news. In South Asian languages, unavailability of datasets is the major hindrance of evaluating fake news. Therefore, my contribution is the induction of a dataset influenced by English fake news dataset for Urdu news; Corpus of Urdu News for Fake News Detection (COUNFND) which has been derived from digital journalism. It contains 480 documents with 120 claims processed by means of conducting traditional machine learning algorithms and text preprocessing techniques. The corpus holds great importance for the evaluation of fake news detection system specifically for Urdu Language. This thesis discusses different techniques that have been applied on the Urdu language datasets and furthermore achieving classification accuracy of about 73 percent effectively. Later on, a new dataset of fake news in Urdu Language has been introduced by Maaz Amjad (Bend the Truth), which has articles divided into fake and real classification, I have applied different supervised learning models and have achieved an accuracy of 0.89 on random forest using four grams techniques which is greater than the baseline accuracy provided with the dataset.