A Benchmark Corpus and Methods for the Urdu Qur’anic Question-Answering System
No Thumbnail Available
Date
2025
Journal Title
Journal ISSN
Volume Title
Publisher
Library Information Services, COMSATS University Islamabad, Lahore Campus
Abstract
Question Answering System aims to provide answers to a question posed in natural
language. A question can be factoid and non-factoid. With the rapid increase in
technologies, the QA task become one of the most challenging and important tasks of
Natural Language Processing (NLP). In previous studies, the problem of the QA system
for the Holy Quran has been explored for Arabic, English, and Indonesian languages. Urdu
has more than 170 million speakers worldwide yet considered low low-resource language.
However, the problem of the QA system for the Holy Quran has not been explored for the
Urdu language due to the unavailability of the dataset. To overcome this research gap, this
study developed a benchmark corpus of the Holy Quran in the Urdu language. We
developed the corpus by carefully mapping the QRCD corpus which was originally in
Arabic language into Urdu. For the mapping of answers, we take the expertise of data
annotators who have a deep knowledge of the Holy Quran in both Arabic and Urdu
language. We conducted experiments in a state-of-the-art BERT based LLMs. We evaluate
the proposed techniques using Exact Match, partial Reciprocal Rank, and F1 score, and
expert judgment.
Description
Keywords
Dr. Rao Muhammad Adeel Nawab, TECHNOLOGY::Information technology::Computer science, Question-Answering System, SP23