Repository logo
Communities & Collections
All of DSpace
  • English
  • العربية
  • বাংলা
  • Català
  • Čeština
  • Deutsch
  • Ελληνικά
  • Español
  • Suomi
  • Français
  • Gàidhlig
  • हिंदी
  • Magyar
  • Italiano
  • Қазақ
  • Latviešu
  • Nederlands
  • Polski
  • Português
  • Português do Brasil
  • Srpski (lat)
  • Српски
  • Svenska
  • Türkçe
  • Yкраї́нська
  • Tiếng Việt
Log In
New user? Click here to register.Have you forgotten your password?
  1. Home
  2. Browse by Author

Browsing by Author "Sana Akhtar Naseer"

Filter results by typing the first few letters
Now showing 1 - 1 of 1
  • Results Per Page
  • Sort Options
  • No Thumbnail Available
    Item
    CSD109_Celebrity Profiling using Content based techniques
    (Library Information Services, COMSATS University Islamabad, Lahore Campus., 2020-11-20) Sana Akhtar Naseer; LHR TP 6211
    Celebrities are persons who get fame due to different work in a different domain (sports, artist, performer, creator, politics, manager, science, professional, and religious). Every celebrity is a well-educated and popular user of social media and provided a lot of writing material and personal detail on it. Nowadays, everyone is interested in celebrities’ lifestyles, but no one can live with a celebrity so, they follow their social media accounts to get their interest and personal detail. Due to the popularity of celebrities, their posting materials are also popular and used by data analysts and researchers. Data analyst and researcher use their text for identification of different demographic traits (gender, age, occupation, birth-year, and fame) which has various applications in forensic, security, marketing, fake profiles identification on online social networking sites, capturing sender of harassing messages, etc. Author profiling is the identification of different demographic traits by analysing of author’s written text. So, Celebrity profiling is a sub-domain of author profiling. In Celebrity profiling there are many existing corpora for English and other European languages but not for South Asian languages, like Urdu or Roman Urdu. Roman Urdu is mostly used by large no. of talkers around the world especially in Face book comments, Twitter tweets, SMS messaging, blogs and chat blogs, etc. So, we aim to develop a corpus on celebrity profiling using Twitter. We are developing a multilingual (Roman Urdu and English) corpus for Asian countries by using content-based features (word and character N-gram) for age, gender, and occupation identification. On the training of the corpus, we manually fetch the username of the verified twitter account of celebrities and list down in Excel sheet. After making list we pass this list to scrapper (Twint, Python Library for scrapping tweets) for scrapping the tweets from twitter. After making a corpus we pre-process tweets (remove URLs, hashtags, mentions, stop words, digits, emojis, and punctuations). After this, we annotate data and apply features Word gram (bi-gram, trigram, tetra-gram, pentagram), Character gram, and content-based techniques for the training of the model. During testing model show accuracy for each predicting class.

DSpace software copyright © 2002-2026 LYRASIS

  • Privacy policy
  • End User Agreement
  • Send Feedback
Repository logo COAR Notify