Repository logo
Communities & Collections
All of DSpace
  • English
  • العربية
  • বাংলা
  • Català
  • Čeština
  • Deutsch
  • Ελληνικά
  • Español
  • Suomi
  • Français
  • Gàidhlig
  • हिंदी
  • Magyar
  • Italiano
  • Қазақ
  • Latviešu
  • Nederlands
  • Polski
  • Português
  • Português do Brasil
  • Srpski (lat)
  • Српски
  • Svenska
  • Türkçe
  • Yкраї́нська
  • Tiếng Việt
Log In
New user? Click here to register.Have you forgotten your password?
  1. Home
  2. Browse by Author

Browsing by Author "Madiha Rasool"

Filter results by typing the first few letters
Now showing 1 - 1 of 1
  • Results Per Page
  • Sort Options
  • No Thumbnail Available
    Item
    Multi-label Document Classification for Urdu Language
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2019) Madiha Rasool; FA17-RCS-010; LHR TP 5779; Dr. Muhammad Salman Khan
    Multi-label document classification is a conventional supervised machine learning task where each document is associated with multiple labels to correctly categorize the meaning of context. All natural languages present this computational linguistic task for a wide series of research domains, such as Natural Language processing, document categorization and text mining. Many benchmark corpora have been developed for multi-label classification task, but most are for English and other European languages. However, there is a dearth of annotated benchmark corpus and methods for the multi-label document classification task for the Urdu language. To fill this gap, we proposed a large benchmark corpus and methods for the multi-label classification of Urdu documents. The proposed corpus contains 600 documents from the field of journalism in the following domains: sports, national, foreign, showbiz and business. The proposed corpus has been manually annotated with USAS (UCREL Semantic Analysis System) scheme which provides 21 major labels and 232 sub labels for fine grained annotation. Each document in the proposed dataset was annotated with minimum two labels and maximum six labels to present a detailed analysis of the context as a supervised multi-label document classification task. To demonstrate how our proposed dataset can be used for the development and evaluation of Urdu multi-label document classification system, three different techniques were applied: (1) Multi-Word Expression (MWE) N-gram (2) Single-Word Expression (SWE) N-gram and (3) Most Frequent Word (MFW) N-gram. Seven different multi-label classifiers were applied on features extracted using three MWE, SWE and MFW methods. Evaluation was carried out using accuracy, exact match, hamming loss and F1 measures. Results show that the (F1 is 0.545) gives best results by using MFW technique on CC and MCC classifier

DSpace software copyright © 2002-2026 LYRASIS

  • Privacy policy
  • End User Agreement
  • Send Feedback
Repository logo COAR Notify