Multi-Label Toxic Comment Classification of Urdu Language
No Thumbnail Available
Date
2020
Journal Title
Journal ISSN
Volume Title
Publisher
Library Information Services, COMSATS University Islamabad, Lahore Campus
Abstract
We are living in a time of technology where a huge amount of information is produced on daily basis on social media websites as it becomes a source to express their views and share ideas with other peoples it also becomes a place for abusive language, personal attacks, and hateful comments. Determining the nature of the comment is difficult and takes a lot of time. Automating the process of detecting toxicity in online comments is the best way to increase user safety and improve online discussions. In this paper, we have produced our dataset of the Urdu language having 20k comments which have been annotated by NLP experts with the following categories: toxic, severe toxic, obscene, threat, insult, and identity hate. The dataset is trained using different machine learning algorithms to find out which model is better in the classification of multi-label toxic comments. Results show that Binary Relevance is the best algorithm in determining the toxicity of comments of the Urdu Language
Description
Keywords
Dr. Waqas Anwar, FA18, Department of Computer Science, Computer Science, Urdu Language, r abusive language, personal attacks