M.Phil / MS

Permanent URI for this collectionhttps://repository.cuilahore.edu.pk/handle/123456789/36

This collection archives the complete set of theses produced by students of the COMSATS University Islamabad, Lahore Campus.

Browse

Search Results

Now showing 1 - 6 of 6
  • Item
    Repeated Buyer Prediction: A Study of Repurchasing Intention of Buyer in E Commerce
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2022) Muhammad Usman; SP20-RCS-014; Dr. Hamid Turab Mirza; LHR TP 8358
    During promotions, retailers frequently obtain a huge number of new customers. However, several of the purchasers attracted are one-time hunters, and the deals are unlikely to have a long-term influence on sales. It was critical for merchants to discover who may be turned into regular loyal consumers and afterward target them to reduce promotion costs and boost return on investment (ROI). It was critical for merchants to discover who could be converted into repeat customers to solve this problem. Merchants may significantly cut promotion costs and increase the return on that investment by focusing on these prospective loyal consumers (ROI). Consumer targeting in the area of internet advertising was generally known to be difficult, especially for first-time consumers. In this work, collect a collection of merchants as well as their associated new buyers gained during the "Double 11" day offer using Tmall.com's long-term user behavior record. This experiment objective was to predict whether new clients would become loyal consumers in the future for certain merchants. In other words, this experiment must estimate the probability, which these new purchasers will buying within the similar merchants again for the next six months. This work suggested employing enhanced merged models (XGBoost as well as LightGBM and Histogram-based gradient boosting machine to forecast a repeat customer and feature engineering through extracting feasible features by which important components would be derived to train the model to prophesy the repeated buyer. These experimental findings suggest that when compared to the original models, this work-combined model may achieve significant performance increases. XGBoost accuracy was 95.97 and AUC was 0.9753. The accuracy of Light GBM was 92.58% and the AUC was 96.15%. Histogram-based gradient boosting machine training accuracy was 96.7% and the testing accuracy was 96.27%.
  • Item
    A Deep Learning Based Prediction of Stock Market Trend using Social Medi
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2022) Aroma Javed; FA20-RCS-006/; LHR TP 8339; Dr. Hamid Turab Mirza
    Machine learning and deep learning are becoming more and more effective techniques for evaluating financial data, encompassing textual, statistical, and digital information. Future stock prediction is a prominent and challenging deep learning topic in the industry. The difficulty in predicting future stock market stems from too many diverse elements that simultaneously influence the amplitude and frequency of stock market rise and falls. In this research work, the main focus is on the problem of stock market trends predictions using social media as a tool. Digital networks are a fast-growing area of information on the Internet. Perhaps one of the most important features is the instant availability of more knowledge and the users' ability to converse swiftly. Different Deep Learning algorithms (like CNN, RNN, GRU, and Bi Directional RNN) were used to forecast stock market trends based on information from social media, as this data might influence investor behavior. Algorithms were used to investigate the impact of social media accounts on stock market prediction performance. The dataset chosen was an expert and public Twitter post from two prominent technology firms, Alphabet Inc. (Google) and Apple Inc, and news data related to these famous firms. The thesis employed deep learning methods, a pre trained language model for economic sentiment analysis, to extract sentiments from tweets. With the help of this research, it will become easy for an investor to invest his money in companies whose stock market values are high on the basis of sentiment classification and will not lead them to any financial crises. SMP aims to anticipate how the stock value of an economic trade will fluctuate in the foreseeable. If shareholders can precisely estimate stock market progression, investors will indeed be able to turn a profit. Finally, the study predicted the trends by modeling the Data on the proposed GRU model, which outperforms the result of other algorithms. The GRU model has shown significant results with an accuracy of 82.41%.
  • Item
    A Study of Crime Analysis and Prediction using Data Mining Techniques
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2022) Syed Huzaifa Ansar; FA19-RCS-030; LHR TP 8338; Dr. Hamid Turab Mirza
    Crime is considered an offense in the social or moral values of a society, due to the constant increase in crimes the safety and security of people in the world are at high risk, because of the severity of this problem, crime analysis is one of the most highlighted research topics from the past few years. In criminology, data mining plays a major role as data mining is a way to discover hidden patterns among the dataset. Researchers are trying to get useful insights for crime prediction by training different machine learning models from real-world data, although a lot of efforts have been made in this context. But most of the models cannot still predict the crime and crime variables such as crime location, date and time. This process becomes even more complex and time-consuming when trying to predict the crime variables. In short, the main goal of this study is to narrow down this gap by developing a system for crime analysis and prediction by using state-of-the-art data mining techniques. In this study, three different deep learning models including Long Short-Term Memory (LSTM), Convolutional Neural Network (CNN) with LSTM layers hybrid model, and Bidirectional Long Short Term Memory (Bi-LSTM) are used to perform a time-series analysis of Chicago Crime dataset from 2001 to 2022. The District-wise time series analysis is performed on the number of crimes for a Month, Week, Day, and night. Mean Square Error (RMSE) and Mean Absolute Error (MAE) are used as evaluation measures to evaluate the performance of models. After the experimentation, results showed that the BILSTM model obtained the best results for the District wise predictions of crimes on the Chicago dataset. The results indicated that the BILSTM model gives the highest performance results (MSE = 7.16, MAE = 2.01) and outperformed all other models in forecasting crimes by day.
  • Item
    Predicting Tourist Destinations based on Interests and Travel Backgrounds using Text Analytics
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2022) ANAM KHURSHID; FA20-RCS-001; LHR TP 8050; Dr. Hamid Turab Mirza
    Tourist destination prediction has arisen as a new topic within text analytics because of the growing adoption of social media. Tourist data is always beneficial to tourism management since it allows them to give personalized services, products, and destinations to future guests. In predicting tourist destinations, machine learning plays a significant role as it is a way to discover hidden patterns among the dataset; researchers have attempted to get valuable understandings by using data from the real world to train models. Many efforts have been made in this regard; however, most of the models still cannot precisely forecast the tourist’s preferences for destinations. In short, the purpose of this study is to narrow down this gap by introducing a system that uses standard machine learning algorithms to forecast destinations chosen by tourists according to their interests and travel histories. This study extracts a dataset from the CouchSurfing.com website to identify the correlation between tourists' interests and travel backgrounds. The data consists of 9575 records of Pakistan users' profile information. To develop a system with standard accuracy, destinations were divided into 8 regions using Google, and interests were divided into 11 categories using the Yahoo category. Using K-means, clustering has been used several times in this study. This research aims to use classification algorithms to cluster the dataset to identify which interests relate to which countries or regions including KNN, Random Forest, AdaBoost, Gradient Boosting, and LDA. After a comparative results analysis of machine learning algorithms, this research found Gradient Boosting has performed best in classification with the highest accuracy of 99.74% and a kappa Score is 0.99. In the end, the system has successfully predicted tourist destinations related to their interests and past travel experiences. The visitors’ interests according to their destination selections have hardly been studied. This research takes a unique strategy from previous studies in that it focuses on this relationship. For text analysis and future travel destination data, the system has also shown good accuracy and surprising results. As a result, it may help travel companies to build marketing plans for tourists with particular interest categories by offering and advertising places.
  • Item
    A Study on Diversification of Online Product Reviews
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2020) Abid Ali; FA17-­RCS-­034; LHR TP 5982; Dr. Hamid Turab Mirza
    Recent studies have boosted the e¬commerce industry which has resulted in increased significance of online product reviews. However, this usefulness of product reviews has also attracted the people who try to manipulate overall product perception by generating fake reviews. Another challenge due to boost in e¬commerce is the information overload which is caused by generation of huge reviews data. This study paves a complete pathway by presenting techniques for removal of spam reviews and by proposing a novel algorithm to retrieve a diversified subset of reviews to reduce the burden of information overload. A diversified set of reviews attempts to cover maximum features of the selected product within a limited number of reviews that ultimately leads to reduction in decision time as well as enhances the credibility and reliability for the user. Spam detection techniques were formulated based on deep learning models whereas novel SENTIMENT AND FEATURE ORIENTED DIVERSIFICATION (SeFOD) algorithm was constructed on the features addressed in each review and the sentiments of the review separately. The proposed models showed prominent results and achieved a maximum spam accuracy of 95.78%, 96.38% and 96.18% for LSTM, GRU and CNN models re spectively. The same results were validated on Yelp hotel reviews dataset. Whereas a new measure for calculating the diversity of the reviews set was adopted named as DivScore. The score nearer to 0 means there is no diversity in the set and hence all the retrieved reviews contain similar features. The far this score goes from 0, the more diversity exists in the diversified set. A DivScore of 7.14 was achieved for the selected product from Daraz reviews dataset while 10.88 was the score when a product was diversified from Yelp reviews dataset. This study can be used by e¬commerce industry to maximize their profits as well as is equally relevant for the general users to better choose relevant product for them
  • Item
    User Interest Base Product Ranking Using Opinions’ Aspects
    (Library Information Services COMSATS University Islamaabad Lahore Campus, 2020) Ahtsham Shoukat; FA17-RCS-011; LHR TP 6401; Dr. Hamid Turab Mirza
    Due to Online shopping trend, a vast number of reviews are generated on daily basis. These reviews help users in selection of products or services. Most of the existing works utilize these reviews in different research areas. Ranking products using these reviews is one of them. Recent studies in the area of research did not focus ranking products on the basis of user specification needs. Previous studies did not focus on ranking products on the basis of user specification. Most of the times the negative polarity of aspects given by users in their reviews is neglected while ranking the products. The main contribution of this study is to take negative polarity of the aspects into consideration, to improve the ranking task. To rank products on the basis of user specification, this research considered both polarities (either positive or negative) of the aspects. In which firstly review is broken down into the sentences to cover the cases in which a review contains multiple senti mentsandvariousaspects. Thereafter aspects are extracted from the reviews of products given in the dataset and polarity of products is calculated on sentence level. Then the polarity is mapped with the aspects extracted from the specific sentence of the review. By using the weightage of the positive and negative polarities products are ranked ac cording to user specification needs. Experimental results show that users now possess the power change the weight of polarities and the resultant ranked list will be according to the user. It means if a user doesnotintendtoincludeaspecificpolarity, hemayignorethesamebeassigningweight 0 to that. This study helps users in selection of good products or services. The instant work is performed for English language reviews. This research opens a gate to the natives of different countries to enhance the work for their native languages
2025 @ COMSATS University Islamabad, Lahore Campus. All rights reserved.