Fake News Detection for Urdu language using Machine Learning Techniques

No Thumbnail Available

Date

2021

Journal Title

Journal ISSN

Volume Title

Publisher

Library Information Services, COMSATS University Islamabad, Lahore Campus

Abstract

The main target of this research is to detect fake news in Urdu language. Although many studies have already been conducted on fake news detection for different languages but not much work has been done for Urdu language and there is a research gap for fake news detection in Urdu language. To compare, evaluate and develop a fake news detection system for Urdu language, standardized and high-quality benchmark resources are needed but unfortunately corpora for Urdu language are not available, even though Urdu is a morphologically rich language and has more than 300 million speakers all over the world. To compensate for this gap, we develop a novel high quality, and a large-sized benchmark corpus for Urdu fake news detection. Only one benchmark dataset is previously available for fake news detection tasks in the Urdu language that contains a total of 900 news articles in which 500 are real news and 400 are fake news articles. Our novel corpus contains 1800 news articles out of which 1000 are real news articles and 800 are fake news articles. Our dataset contains five domain data (Business, Health, Showbiz, Sports, and Technology). These news articles are tagged manually into real and fake news tags. These news articles are collected from different authentic news sources. After corpus creation, feature extraction is performed using different combinations of word n-gram, character n-gram, and function n-gram. The range of this n-gram varies from 1 to 6 for our experiments. We use different weighting schemes such as Term Frequency-Inverse Document Frequency (TF-IDF), Log entropy, Binary weighting schemes, etc. After extracting features from the text, state-of-the-art machine learning techniques are applied to the corpus to evaluate the fake news detection system. In machine learning techniques, Multinomial Naïve Bayes (MNB), Linear Support Vector Classifier, Logistic regression, Random Forest, Decision Tree, Bernoulli Naïve Bayes (BNB) and Ada-Boost are used. Each classifier is trained and tested separately. Standard evaluation measures are used to check the working of the fake news detection system i.e., accuracy, precision, recall, F1-score. We compare our results with previous research results and Linear SVC and Logistic Regression outperforms other classifiers.

Description

Keywords

Dr. Hasan Jamal, fa18, Department of Computer Science, TECHNOLOGY::Information technology::Computer science, Fake News Detection, Urdu language, Machine Learning Techniques

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By