CSD109_Celebrity Profiling using Content based techniques

No Thumbnail Available

Date

2020-11-20

Journal Title

Journal ISSN

Volume Title

Publisher

Library Information Services, COMSATS University Islamabad, Lahore Campus.

Abstract

Celebrities are persons who get fame due to different work in a different domain (sports, artist, performer, creator, politics, manager, science, professional, and religious). Every celebrity is a well-educated and popular user of social media and provided a lot of writing material and personal detail on it. Nowadays, everyone is interested in celebrities’ lifestyles, but no one can live with a celebrity so, they follow their social media accounts to get their interest and personal detail. Due to the popularity of celebrities, their posting materials are also popular and used by data analysts and researchers. Data analyst and researcher use their text for identification of different demographic traits (gender, age, occupation, birth-year, and fame) which has various applications in forensic, security, marketing, fake profiles identification on online social networking sites, capturing sender of harassing messages, etc. Author profiling is the identification of different demographic traits by analysing of author’s written text. So, Celebrity profiling is a sub-domain of author profiling. In Celebrity profiling there are many existing corpora for English and other European languages but not for South Asian languages, like Urdu or Roman Urdu. Roman Urdu is mostly used by large no. of talkers around the world especially in Face book comments, Twitter tweets, SMS messaging, blogs and chat blogs, etc. So, we aim to develop a corpus on celebrity profiling using Twitter. We are developing a multilingual (Roman Urdu and English) corpus for Asian countries by using content-based features (word and character N-gram) for age, gender, and occupation identification. On the training of the corpus, we manually fetch the username of the verified twitter account of celebrities and list down in Excel sheet. After making list we pass this list to scrapper (Twint, Python Library for scrapping tweets) for scrapping the tweets from twitter. After making a corpus we pre-process tweets (remove URLs, hashtags, mentions, stop words, digits, emojis, and punctuations). After this, we annotate data and apply features Word gram (bi-gram, trigram, tetra-gram, pentagram), Character gram, and content-based techniques for the training of the model. During testing model show accuracy for each predicting class.

Description

Keywords

CSD109_Celebrity Profiling using Content based techniques, Computer science, FA16, SP16

Citation

Endorsement

Review

Supplemented By

Referenced By