Cross-Lingual Paraphrase Detection for English Urdu Language Pa

dc.contributor.authorNida Waheed
dc.contributor.authorSP19-RCS-025
dc.contributor.authorLHR TP 7598
dc.contributor.authorDr. Rao Muhammad Adeel Nawab
dc.date.accessioned2026-02-16T04:51:55Z
dc.date.issued2021
dc.description.abstractThe main aim of the Cross-Lingual Paraphrase Detection (CLPD) task is to identify whether a text pair (text 01 in language A and text 02 in language B) are paraphrased of one another or not. With the massive increase in technological advancements in digital media and internet facilities, the availability of digital content across multiple languages has become free as well as easy to access. Besides paraphrasing, there has been a rise in cross-lingual paraphrasing since the development of advanced text processing tools. Moreover, the facility of translating existing texts from the original language into any other desired language using freely available translation tools has worsened the problems and triggered an increase in cross-lingual paraphrasing. Previously, there has been plenty of work done for the cross-lingual paraphrase detection in different languages but the downside is that under-resourced languages are still neglected. Urdu is a widely spoken language, despite the fact, the work done for the Urdu language in the CLPD domain is notably less because of the complexity and challenges of the Urdu language. It still comes in the category of low-resource languages. The research in the text analysis domain is mostly focused on well-resourced languages such as English and other European languages. The CLPD problem has formerly been explored for English-Czech, English German, English-Spanish, English Chinese, etc. language pairs. However, the CLPD problem has not yet been studied and reported for the English-Urdu language pair. Although, Cross-Lingual Paraphrasing is now being used enormously in various domains such as content writing, blog posts, and academics, especially in research works related to different languages. To overcome this research gap, a benchmark corpus and approaches are presented through this study for CLPD. The first objective of this research work is to develop a large benchmark Cross-Lingual Paraphrase Detection Corpus (CLPD-EU) for English-Urdu language pair at the sentence level by using a semi-automatic translation approach. For the creation of the CLPD-EU corpus, an existing corpus named Microsoft Research Paraphrase x Corpus (MRPC) is used. The second chief objective of this research work is to make a thorough comparison of classical machine learning (Bilingual Dictionary Based Approaches), Cross-lingual Word Embedding Based Approaches, Transfer Learning (Sentence Transformers Based Approaches) Approaches, and Feature Fusion Approaches on our proposed corpus. The results exhibited that our proposed Feature Fusion Approach 1 has shown notable performance for CLPD. The evaluation of all these approaches is carried out using the weighted average of Precision, Recall, and F1 scores.
dc.identifier.urihttps://repository.cuilahore.edu.pk/handle/123456789/1653
dc.language.isoen
dc.publisherLibrary Information Services, COMSATS University Islamabad, Lahore Campus
dc.relation.ispartofseriesLHR TP 7598
dc.subjectDr. Rao Muhammad Adeel Nawab
dc.subjectSp19
dc.subjectDepartment of Computer Science
dc.subjectTECHNOLOGY::Information technology::Computer science
dc.subjectCLPD
dc.subjectEnglish-Urdu Language Pair
dc.titleCross-Lingual Paraphrase Detection for English Urdu Language Pa
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
SP19-RCS-025 - Thesis.pdf
Size:
4.19 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
319 B
Format:
Item-specific license agreed to upon submission
Description:

Collections