Developing a Gold Standard Benchmark Corpus and Methods for Multi-label Emotion Classification for Urdu Language

dc.contributor.authorAzqa Abdul Ghani
dc.contributor.authorSP22-RCS-025
dc.contributor.authorLHR TP 8685
dc.contributor.authorDr. Rao Muhammad Adeel Nawab
dc.date.accessioned2026-02-16T10:56:12Z
dc.date.issued2023
dc.description.abstractEmotion Classification (EC) involves detecting the emotional tone conveyed by text, while Multi-Label Emotion Classification (MLEC) aims to assign relevant emotion labels that accurately reflect the author's state of mind. This research focuses on MLEC in Urdu text. In previous studies, only one work has been done for the problem of MLEC, including the development of a dataset based on Urdu Tweets using six emotions [5]. However, the problem of MLEC has not been explored for 12 emotions and diverse genres. To fulfill this research gap, this research study has developed a large benchmark corpus of 16,011 Urdu reviews, manually annotated with 12 emotions and gathered from 9 different genres. Furthermore, this research study has Developed, Applied, Evaluated, and Compared Classical Machine Learning, Deep Learning and Transfer Learning based methods. Whereas in Transfer Learning methods, we have implied 6 distinct sentence transformers (xlm-r-100langs-bert-base-nli-stsb-mean-tokens, xlm-r-100langs-bert-base nli-mean-tokens, gtr-t5-large, all-mpnet-base-v2, distiluse-base-multilingual-cased-v2, and LaBSE), Also the Combined Feature-based approach is applied which averages the features of all the columns of the sentence transformer model and then the Machine Learning models are applied. The proposed transformers-based approach outperforms with an F1 score of 0.4562 using the ‘xlm-r-100langs-bert-base-nli-stsb-mean-tokens’ sentence transformer
dc.identifier.urihttps://repository.cuilahore.edu.pk/handle/123456789/1732
dc.language.isoen
dc.publisherLibrary Information Services, COMSATS University Islamabad, Lahore Campus
dc.relation.ispartofseriesLHR TP 8685
dc.subjectsp22
dc.subjectDr. Rao Muhammad Adeel Nawab
dc.subjectDepartment of Computer Science
dc.subjectTECHNOLOGY::Information technology::Computer science
dc.subjectGold Standard Benchmark
dc.subjectMulti-label Emotion
dc.subjectUrdu Language
dc.subjectEmotion Classification (EC)
dc.titleDeveloping a Gold Standard Benchmark Corpus and Methods for Multi-label Emotion Classification for Urdu Language
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
8685.pdf
Size:
1.14 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
319 B
Format:
Item-specific license agreed to upon submission
Description:

Collections