M.Phil / MS

Permanent URI for this collectionhttps://repository.cuilahore.edu.pk/handle/123456789/36

This collection archives the complete set of theses produced by students of the COMSATS University Islamabad, Lahore Campus.

Browse

Search Results

Now showing 1 - 10 of 10
  • Item
    Process Model Generation from Textual Descriptions
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2021) Aleena Nazir; FA18-RCS-014; LHR TP 7290; Dr. Abid Sohail Bhutta
    Process models are an important source to gather the information on organizational workflows and also represent the first point of process analysis and improvement. In organizations, business process modeling is an important tool to understand and automate the business processes. However, in many organizations the existing documentation of business processes is difficult to understand by the analyst. So, in real-life organizations, the high complexity of business processes is constantly raising an issue. Due to this, the establishment of process models for the business process is becoming a challenge for the stakeholders. So, the process model extraction from the business processes may helpful to minimize the process modeling effort. But, manual generation of business process models is a time taking task for the stakeholders. Yet, to assist this task, new methods can be implemented for the automation of process design phase. An approach to generate process models from the textual descriptions has been developed. In this approach, we use different natural language processing techniques to define a set of mapping rules. Through these mapping rules, we extract the elements business process models from textual descriptions and also checked the correct sequence of extracted elements for business process models. By combining the business process elements according to their sequence, process model is generated
  • Item
    On Enhancing the Quality of Business Process Model’s Activities Labels
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2021) Urooj Fatima; FA18-RCS-006; LHR TP 7291; Dr. Abid Sohail Bhutta
    Business process models describe their operations, events, and control flows through graphical illustrations to increase the knowledge and awareness of business processes. Large corporations use models to document and design business processes. With the growing number of business process models and trained modelers, modeling initiatives demand quality assurance. Nevertheless, checking the quality of the process model, especially its activity labels, is a challenge. In labels, synonymy, vagueness, homonymy, incorrect labeling, as well as different modeling styles result in ambiguity, uncertainty, and misunderstandings. Quality of activity labels rely on precise and fitful words which are according to the domain process models taking quality parameters under consideration. The problem arises when the activity labels are too short and provide limited information and words facing the zero-derivation problem. For this purpose, algorithms have been deployed which will recognize, identify and check the labeling styles of a process model. Activity labels has been extracted automatically. Further, NLP techniques like WordNet has been used for the analysis of activity labels. In this study, the quality of textual labels in activities of a process models is addressed. Activity labels has been analyzed using a collection of business process models based on medical chronic diseases. Results obtained by deployment of algorithms on automatically extracted labels confirms the applicability and accuracy of proposed techniques.
  • Item
    A Resource Efficient Method for Indexing Hidden Web Using Rank Based Web Crawling Techniques
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2021) Abdul Mannan; FA18-RCS-026; LHR TP 7282; Dr. Muhammad Hasanain Chaudary
    Today we are seeing a shift in understanding and behavior of individuals toward anonymity and privacy. As a result, not only usage of virtual private network is increasing, but also more and more people are converging towards hidden web. Hidden web is a server less chain-based architecture, which can only be accessed using specific proxies and gateways and provides anonymity-using chain of interconnected nodes over public IP so that even if node is compromised anonymity of user is maintained. This form of security has its own drawbacks as overall network speed depends on node with lowest connectivity. Subsequently, crawling becomes costly since speed is proportional to the latency of generated Tor circuit. Another challenge in crawling Hidden web is the volatile nature of hosted services provided on it. Due to its anonymous nature, illegal services are prevalent. Consequently, this network is highly monitored becoming hot bed of banned websites some of which become live on new nodes while other stay down. This causes loss of time and resources crawler used to mine those dead URL. We want to propose a crawler that can mine data in low latency Tor network, auto tuning its configuration according to the state of the network and giving rank to websites in line with their content. This rank would be used to calculate crawling depth of a specific service at a given time, further improving as service stays alive
  • Item
    Class-based Unsupervised Parts-of-Speech Tagging for Low Resource Languages
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2020) Zunaira Ashraf; FA18-RCS-011; LHR TP 6411
    The Unsupervised Parts-of-Speech Tagging is a challenging task for linguistic structures especially for low-resource languages. Most of the recent research on this topic used Hidden Markov Model (HMM) with different variations in models, datasets, and hyper-parameter values. Many of them worked upon high-resource mono, bi, or multi-lingual, in cross-lingual compared the dictionaries of two languages and then assigned tags to latent words, but there isn’t any proper or efficient model learned upon low-resource languages especially for unsupervised learning. Majority of them tuned the hyper-parameters after each iteration but few learned the parameters and fixed these values for all iterations of all languages. We focus on the main concept of Unsupervised POS Tagging i.e., clustering the unannotated data through grouping of similar word-tags by using standardized tag sets/clusters and work on Urdu low-resource languages and on English too. We designed a creative model which followed the common approach, that is to maximize the probability of hidden structure under observed data. Model parameter’s Maximum-likelihood estimation (MLE) is used for this purpose. Bayesian HMM improved performance in part-of-speech tagging for unsupervised learning. Bayesian approach performed over all possible hyper-parameters instead of single set of parameter values. Then, choose that set/pair of hyper-parameter values which gives best results. Our model has the structure of regular trigram HMM, as its accuracy for English language is closer to the state-of-the-art unsupervised selective model. And is 9% greater than the maximum-likelihood estimation. But for Urdu language, there is not any work in unsupervised POS tagging, so its accuracy is not compared with any model I.e., 77.6%, it’s less due to missing information of Urdu language. In future, we are going to improve our model until get its maximum accuracy.
  • Item
    Requirement Change Prediction Model For Small Software systems
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2020) Rida Fatima; FA18-RCS-025; LHR TP 6416; Dr. Farrukh Zeshan
    The software industry is rising gradually, it seems that there is no end to software development. Modern era software projects are complex and consist of many components, so change is unavoidable in these projects. Therefore, change in the software requirements document must be predicted at an early stage to preserve the resources. There are a lot of requirement change prediction models proposed in the literature and they are addressing a smaller number of variables and including experts knowledge for analyzing the requirement specification document. In this work, we have focused on the small software systems which have focused few variables of requirement change prediction model.The variables have a major effect on the software requirements and these are analyzed by the stakeholders, developers and experts with the questionnaire method. Their knowledge is incorporated in the Bayesian network as conditional probabilities of independent and dependent variables.This is actually the proposed model for software requirements change prediction.We used an algorithm with the model by utilizing variable elimination method to obtain the posterior probability of the revisions in software requirement document. However this model can also be used for the large software systems because it is effecent to resolve their problems also. the The proposed model is evaluated by hypothesis testing, sensitivity analysis and by comparing with the existing models. The results obtained proved to be promising as it decreases the probability of revisions in the requirement document. We obtained the 0.42 probability for the revisions in the requirement document so the results have improved the existing model predictions
  • Item
    Brain-controlled Intelligent Mobility System for Neurally Damaged Quadriplegic Patients
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2020) Syed Abu Huraira Hussain; FA18-RCS-022; LHR TP 6415; Prof. Dr Syed Asad Hussain
    Brain-Computer Interface (BCI) provides direct interaction between the brain and computing devices without the need for any physical contact, which makes it a useful tool in applications designed for paralytic patients. Power Wheelchair (PW) control through BCI is one of the principal applications for paralytic patients. The practical implementation of brain-controlled PW is still questionable due to the system’s performance issues, the patient’s physical state, and behavioral sensitivity. Over time, patient’s physical and mental state changes which affect the brain signals and could alter the results of the system. So, there is a challenge to establish a system that is accurate, safe, and efficient enough to act in real-time with a greater number of control commands. To increase the adaptability of BCI based PW, a system that not only maps the brain signals to control commands but also provides the contingency mechanism and speed control is presented. The contingency mechanism of the system helps to tackle such stressful situations by switching the control to joystick and in case of a medical emergency as well. Moreover, the system continuously monitors the mental state along with feature extraction from the brain signals. This continuous monitoring of the mental state will trigger the contingency mechanism when needed, which complements an additional layer of safety and makes the system more compliant. The presented system involves user interference only for the brain signal acquisition after that it performs all the tasks on the computer along with the microcontroller smartly on its own. To test the systems' credibility and efficiency; a series of experiments were performed. The results show that the system is reliable and safe enough. The system also shows satisfactorily True Positive Rate (TPR) and False Positive Rate (FPR) with an average time of 7.7 seconds to generate the interpretable brain signal from the user. To monitor the mental state an auditory-based experiment has been designed which allows the system to learn about the personalized mental states of the user. The system predicts the mental state of the user with an average accuracy of 74.34%
  • Item
    Call Admission and Channel Assignment Framework Ensuring QoS for Multimedia Cellular Cognitive Radio Networks
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2020) Athar Ikhlaq; FA18-RCS-029; LHR TP 6417; Dr. Amjad Ali
    In this astonishingly growing technological era, wireless multimedia applications are increasing exponentially based on these technological advancements. Therefore, wireless multimedia traffic needs a spectrum for transmission. Radio Spectrum is Fixed and used statically worldwide. Cognitive radio in wireless multimedia networks has gained popularity over the past few years because of reducing the scarcity barrier by sensing the radio spectrum. This mechanism provides opportunities to Secondary Users (SUs) to utilize the unused bandwidth of the Primary Users (PUs) licensed spectrum over the Cellular Cognitive Radio Networks (CCRNs). However, multimedia applications demand stringent Quality of Service (QoS) for seamless Communication. The unused radio spectrum is dynamically available with location and time-dependencies. These dependencies starkly affect channel reliability (i.e., the channel's ability to serve the particular SU) in terms of QoS parameters. In this research, we have proposed a non-cooperative game theory-based Call Admission framework that is allocating resources to the admitted calls based on their desired required quality of services. For this purpose, the mechanism does appropriate channel assignment as per the SUs request for efficient resource utilization ensuring seamless multimedia transmissions over WCCRNs. The algorithm is providing a suitable channel selection for the maximum QoS-based spectrum utilization. The system guarantees the QoS defined parameters for multimedia transmissions so that the appropriate channel will be assigned to the SU based on its desired QoS requirements. Due to the opportunistic behavior, i.e., the dynamicity in the availability of the resources, MSU has to utilize the best suitable and available resource at the time of request entertainment. This behavior of the radio spectrum availability is due to the PU priority and we have tackled it by applying the non-cooperative mechanism. For each request, 4 actions have been provided by this theory, either the request will have what it demands or it will be provided with a greater one to ensure quality. If these 2 cases failed in entertaining the request, then the system will check for any available resource with a lower capacity but the satisfactory quality of services level in 3rd case, and assign it to the enrolled request. If none of these cases is possible then the enrolled request is simply discarded in case 4, as the system has no resources to entertain the respective request. Once a channel is assigned to a specific request against any of 3 cases, the system will start checking for primary user arrival. if detected, the allocated channel is made vacant and the transmission gets suspended. If primary user interference is not detected then the request will continuously transmit data until completion. The suspended request is entertained at priority next time to reduce the handoff rate. We have compared our results with the already available scheme in which they have used the bidding mechanism to allocate resources. For the proposed system evaluation, we have used the fitness function to elaborate the results gathered by the proposed system in terms of entertaining the requests as per their demand. The simulation environment has shown the results that the proposed system has outperformed the bidding technique to entertain more QoS-based requests.
  • Item
    Multi-Label Toxic Comment Classification of Urdu Language
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2020) Mian Ahmed Shafiq; FA18-RCS-002; LHR TP 6408; Dr. Waqas Anwar
    We are living in a time of technology where a huge amount of information is produced on daily basis on social media websites as it becomes a source to express their views and share ideas with other peoples it also becomes a place for abusive language, personal attacks, and hateful comments. Determining the nature of the comment is difficult and takes a lot of time. Automating the process of detecting toxicity in online comments is the best way to increase user safety and improve online discussions. In this paper, we have produced our dataset of the Urdu language having 20k comments which have been annotated by NLP experts with the following categories: toxic, severe toxic, obscene, threat, insult, and identity hate. The dataset is trained using different machine learning algorithms to find out which model is better in the classification of multi-label toxic comments. Results show that Binary Relevance is the best algorithm in determining the toxicity of comments of the Urdu Language
  • Item
    Lexical Sample Word Sense Disambiguation for Urdu Language – Data and Methods
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2020) Hamza Jawed; FA18-RCS-037; LHR TP 5990
    In all natural-languages, there exists a range of words that have multiple meanings. Word Sense Disambiguation (WSD) is the process of automatically determine the correct sense of a word in a given context. Thus, this lexical ambiguity is observed as one of the significant and critical issues in natural language processing (NLP). To build, evaluate and compare WSD methods the standard assessment resources are required. To develop benchmark WSD corpora and methods for a broad variety of languages, a range of initiatives are needed. But despite the fact, that there are more than 325 million speakers of Urdu and a larger volume of the digital text of Urdu is available online, the standard is deficient, high-quality big corpus for the WSD task. To overcome that gap, the current work aims to create a novel and large size benchmark corpus for the task of Urdu Lexical Sample WSD with the application of classical machine learning and deep learning approaches on this large corpus. The work comprises manually crafting the 500 most frequent polysemic words with 127 General Urdu terms, 136 terms from News, 72 terms from Religion, 21 terms from Science, 52 terms from Literature, 29 terms from Medical and 63 terms from Finance from UrMono Corpus (a benchmark corpus of Urdu with 95.4 million Urdu tokens) and Urdu Lughat (a sense inventory and standard dictionary of Urdu Language). After the creation of a corpus, an inspection will be conducted on the ambiguous terms for the verification of the WSD system through four deep learning techniques (1) Multi-Layer Perceptron (MLP), (2) Simple Recurrent Neural Network (RNN), (3) Long-Shot-Term-Memory (LSTM), and (4) Bidirectional Long Short Term Memory (Bi-LSTM). Five classical machine learning algorithms will also be tested on the same dataset. A comparison between of results of classical machine learning and deep learning techniques reflects that machine learning models outperform the deep learning techniques.
  • Item
    Resource Creation and Evaluation of Aspect Based Sentiment Analysis in Urdu
    (Library Information Services COMSATS University Islamaabad Lahore Campus, 2020) LHR TP 6412; Dr. Muhammad Waqas Anwar; FA18-RCS-017; Sadaf Rani
    Sentiment analysis is an emerging research area of Natural Language Processing (NLP). Along with the rise of people generated content on social sites, sentiment analysis has gained more importance. One of the challenging tasks of sentiment analysis is Aspect Based Sentiment Analysis (ABSA). ABSA is a task of identifying the sentiment at aspect level. Right now, mostly researchers have focused on the English language but very less attention on resource poor languages like Urdu, Arabic etc. To the best of our knowledge, there is no publicly available dataset on ABSA in Urdu language. In this study, we focused on the resource creation for the task of ABSA in Urdu language. This task is further divided into four subtasks i.e., Aspect Term, Aspect Term Polarity, Aspect Category and Aspect Category Polarity. Our dataset is comprised of 6672 sentences which are collected form twitter using Twitter API’s. For annotation, we prepared standard guidelines which are according to the SemEval. Evaluation is also performed by extracting n gram features at word level and char level. TF-IDF vectorizer is used to vectorize the data into machine readable form. Then we applied different machine learning algorithms. Experiments are performed by Naïve Bayes (NB), support Vector Machine (SVM), Logistic Regression (LR), Random Forest (RF) and performance is measured by four parameters i.e., accuracy, precision, recall, f1 measure. We achieved 71% accuracy in Aspect Term task, 61% in Aspect Category, 74% in Aspect Term Polarity and 80% in Aspect Category Polarity The evaluation results depict that there is need to do more work in this task. We provided baseline evaluation for researcher community to further compare their results of ABSA systems.
2025 @ COMSATS University Islamabad, Lahore Campus. All rights reserved.