Classification of Non-Functional Requirements by using Machine Learning and Natural Language Processing Techniques

No Thumbnail Available

Date

2022

Journal Title

Journal ISSN

Volume Title

Publisher

Library Information Services, COMSATS University Islamabad, Lahore Campus

Abstract

The software development process consists of a series of phases e.g., requirement engineering, design, coding, and testing, and each phase is critical for fulfilling the needs of a software user. The behavior and development of a software system are restricted by NFRs, which define the general characteristics or attributes the final system must possess. An automated technique for identifying NFRs is necessary to find NFRs for the software development lifecycle. The formatting and classification of NFRs subclasses such as performance efficiency, usability, maintainability, portability, security, and reliability are less discussed and there is a lack of a larger publicly available dataset. In addition, the datasets are not semi-formated to overcome inherent ambiguities and incompleteness in NFRs specified in natural language. NFRs are specified in natural language which faces the challenges of overcoming confusion, redundancy, and ambiguity due to a lack of structure in the requirements specification. The NFRs such as performance efficiency, usability, maintainability, portability, security, reliability, etc., will be the focus of our project. This research will use a larger dataset with over more than 2000 NFRs which were created through processing 315 software requirement specifications (SRS) documents of BS (CS) and BS (SE) final year projects (FYP) of CUI, Lahore, and requirements will be semi-formally specified using Rupp's boilerplate to avoid natural language specification problems. The software requirements were passed through several steps including normalization and feature extraction methods like TF-IDF. Several Machine Learning (ML) and Deep Learning (DL) experiments were carried out, including Logistic Regression (LR), Bernoulli Naive Bayes (BNB), Decision Tree (DT), Support Vector Machine (SVM), Multinomial Nave Bayes (MNB), Random Forest (RF), CNN, and Long Short-Term Memory (LSTM), and Bi-LSTM algorithms for the classification of NFRs. SVM performed well in the classification of NFRs with F1-score of 0.82.

Description

Keywords

Dr.Touseef Tahir, sp21, Computer Science, Department of Computer Science, Non-Functional Requirements, Machine Learning, Natural Language Processing Techniques

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By