Multi-Label Toxic Comment Classification of Urdu Language

No Thumbnail Available

Date

2020

Journal Title

Journal ISSN

Volume Title

Publisher

Library Information Services, COMSATS University Islamabad, Lahore Campus

Abstract

We are living in a time of technology where a huge amount of information is produced on daily basis on social media websites as it becomes a source to express their views and share ideas with other peoples it also becomes a place for abusive language, personal attacks, and hateful comments. Determining the nature of the comment is difficult and takes a lot of time. Automating the process of detecting toxicity in online comments is the best way to increase user safety and improve online discussions. In this paper, we have produced our dataset of the Urdu language having 20k comments which have been annotated by NLP experts with the following categories: toxic, severe toxic, obscene, threat, insult, and identity hate. The dataset is trained using different machine learning algorithms to find out which model is better in the classification of multi-label toxic comments. Results show that Binary Relevance is the best algorithm in determining the toxicity of comments of the Urdu Language

Description

Keywords

Dr. Waqas Anwar, FA18, Department of Computer Science, Computer Science, Urdu Language, r abusive language, personal attacks

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By