Fake News Detection and Corpus Development for Urdu News

dc.contributor.authorMuhammad Ismail Suri
dc.contributor.authorSP17-RCS-009
dc.contributor.authorLHR TP 7287
dc.contributor.authorDr. Muhammad Waqas Anwar
dc.date.accessioned2026-02-13T10:59:07Z
dc.date.issued2021
dc.description.abstractWith the growing technological advancements in the present era, the curiosity to remain updated about the daily happenings in the world has opened up various opportunities for news dissemination. Access to information with ease has made the propagation of news much efficient and cost effective. However, online news is a double-edged sword that creates a room for spreading fake news which intends to mislead and exaggerate the reality. This issue is a rising problem leading towards several real world sociopolitical and cultural impacts. To curb the menace of fake news different organizations are working to handle it, but majority of the work is being done in English, unfortunately the language of Urdu in this dimension remains untouched. This research is focused on detecting fake news in Urdu language. The demand is to have a detecting mechanism that can predict between fake and real news. In South Asian languages, unavailability of datasets is the major hindrance of evaluating fake news. Therefore, my contribution is the induction of a dataset influenced by English fake news dataset for Urdu news; Corpus of Urdu News for Fake News Detection (COUNFND) which has been derived from digital journalism. It contains 480 documents with 120 claims processed by means of conducting traditional machine learning algorithms and text preprocessing techniques. The corpus holds great importance for the evaluation of fake news detection system specifically for Urdu Language. This thesis discusses different techniques that have been applied on the Urdu language datasets and furthermore achieving classification accuracy of about 73 percent effectively. Later on, a new dataset of fake news in Urdu Language has been introduced by Maaz Amjad (Bend the Truth), which has articles divided into fake and real classification, I have applied different supervised learning models and have achieved an accuracy of 0.89 on random forest using four grams techniques which is greater than the baseline accuracy provided with the dataset.
dc.identifier.urihttps://repository.cuilahore.edu.pk/handle/123456789/1562
dc.language.isoen
dc.publisherLibrary Information Services, COMSATS University Islamabad, Lahore Campus
dc.relation.ispartofseriesLHR TP 7287
dc.subjectDr. Muhammad Waqas Anwar
dc.subjectsp17
dc.subjectDepartment of Computer Science
dc.subjectTECHNOLOGY::Information technology::Computer science
dc.subjectFake News Detection
dc.subjectUrdu News
dc.titleFake News Detection and Corpus Development for Urdu News
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
7287.docx
Size:
399.32 KB
Format:
Microsoft Word XML

License bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
319 B
Format:
Item-specific license agreed to upon submission
Description:

Collections