A Benchmark Corpus and Methods for the Urdu Qur’anic Question-Answering System

No Thumbnail Available

Date

2025

Journal Title

Journal ISSN

Volume Title

Publisher

Library Information Services, COMSATS University Islamabad, Lahore Campus

Abstract

Question Answering System aims to provide answers to a question posed in natural language. A question can be factoid and non-factoid. With the rapid increase in technologies, the QA task become one of the most challenging and important tasks of Natural Language Processing (NLP). In previous studies, the problem of the QA system for the Holy Quran has been explored for Arabic, English, and Indonesian languages. Urdu has more than 170 million speakers worldwide yet considered low low-resource language. However, the problem of the QA system for the Holy Quran has not been explored for the Urdu language due to the unavailability of the dataset. To overcome this research gap, this study developed a benchmark corpus of the Holy Quran in the Urdu language. We developed the corpus by carefully mapping the QRCD corpus which was originally in Arabic language into Urdu. For the mapping of answers, we take the expertise of data annotators who have a deep knowledge of the Holy Quran in both Arabic and Urdu language. We conducted experiments in a state-of-the-art BERT based LLMs. We evaluate the proposed techniques using Exact Match, partial Reciprocal Rank, and F1 score, and expert judgment.

Description

Keywords

Dr. Rao Muhammad Adeel Nawab, TECHNOLOGY::Information technology::Computer science, Question-Answering System, SP23

Citation

Collections

Endorsement

Review

Supplemented By

Referenced By