M.Phil / MS
Permanent URI for this collectionhttps://repository.cuilahore.edu.pk/handle/123456789/36
This collection archives the complete set of theses produced by students of the COMSATS University Islamabad, Lahore Campus.
Browse
2 results
Search Results
Item Multi-Aspect Hate Speech Analysis for Roman Urdu Text(Library Information Services, COMSATS University Islamabad, Lahore Campus, 2021) Aniqa Khalid; SP19-RCS-004; LHR TP 7597; Dr. Waqas AnwarWe live in the age of technology where a large amount of information is produced daily on social media sites as it becomes a source for expressing their opinions and sharing ideas with other people, it also becomes a place for abusive language, personal attacks, and hateful comments. Determining the nature of the suspension is difficult and time-consuming. Automating the process of hate speech analysis in online conversations is the best way to ensure user security and improve online conversations. In this study, we have produced our dataset for Roman Urdu containing more than 3k comments which were annotated by NLP experts with the following aspects: Hostility, directness, target and group. The dataset is trained using various deep learning & machine learning algorithms for figuring out which model is the best at classifying multi-aspect hate speech. The results showed that logistic regression and bi-LSTM are the best algorithm in determining the toxicity of Roman Urdu text.Item Toxic Comment Classification of Roman Urdu Text(Library Information Services, COMSATS University Islamabad, Lahore Campus, 2019) Waheed Abbas; FA17-RCS-013; LHR TP 5624; Dr. Rao Muhammad Adeel NawabWe are living in an era of technology where most of us have easy access to the inter net. Due to internet bloom, the use of social media and discussion forum, in particular for communication, increase tremendously in recent years. But this advancement also opens doors for trolls who poison these social media and forums by their abusive behav ior toward others. Detection of toxic language online is becoming an important issue in recent years. A toxic comment is defined as rude, unreasonable, or disrespectful com ments which likely to make you leave a conversation. Current methods for handling toxicity online are often heavily dependent on manual moderation, and not scalable enough to handle growing users day by day. In this thesis, we investigated the problem of toxic language detection for Roman Urdu (transliteration of Urdu in English) text because Roman Urdu is widely used to communicate online in Pakistani and Indian community. Various work has been done to detect toxicity in the English language, but due to unavailability of large Roman Urdu corpora, there is no work done so far in Roman Urdu. The aim of this thesis work is two folds: first is to develop large scale corpus for detection of toxic language in Roman Urdu, second is to apply different proven classical machine learning and state-of-the-art deep learning methods on the developed corpus. In the first step, we collected data from the YouTube comment section to develop Roman Urdu corpus. The final corpus consists of 10,863 unique comments where we have 5,887 toxic and 4,976 non-toxic comments. We performed manual annotation of corpus and got it reviewed from another annotator. We applied proven classical ma chine learning algorithms like Logistic Regression, K-Nearest Neighbors, Naive Bayes, Random Forest, Support Vector Machine and Multi-layer Perceptron along with state of-the-art deep learning methods like Long Short-Term Memory, Bidirectional Long Short-Term Memory, and Convolutional Neural Networks. Best results obtained from Random Forest with character 4-gram (Accuracy = 0.964, Precision = 0.980, Recall = 0.953, F1 = 0.966).