Urdu Sentential Paraphrase Generation using GAN

dc.contributor.authorNimra Razzaq
dc.contributor.authorSP22-RCS-003
dc.contributor.authorLHR TP 8681
dc.date.accessioned2026-02-17T05:05:01Z
dc.date.issued2023
dc.description.abstractParaphrasing can be described as the task of restatement of a piece of text that carries the same meaning but with different expressions. It requires manual effort, domain knowledge, and language expertise to paraphrase a text.In contrast, automatic paraphrase generation involves using Natural Language Processing (NLP) methods to automatically produce a paraphrased version of a text. Automatic generation of paraphrases is not only helpful for content writers or authors but has a range of applications in other NLP tasks, e.g., Information Retrieval (IR), knowledge extraction, summarizing text etc. Previous efforts for automatic paraphrase generation have been largely conducted in the English language. To the best of our knowledge, there exist no corpora for paraphrase generation other than the English language. Moreover, the issue of automatic paraphrase generation using large volumes of data and utilizing more recently proposed models (GAN and mBART) has not been investigated for the Urdu language. To bridge this research gap, the purpose of this research work is twofold: Firstly, we created the first Urdu Sentential Paraphrase Generation Corpus, using manual and semi-automatic approaches. The corpus contains 8,000 sentence pairs mined from existing Urdu corpora, online news sources, or machine-translated from already available English corpora. Secondly, we applied newly proposed state-of-the-art methods (GAN, mBART)to the proposed corpus. The evaluation was carried out automatically using ROUGE and BLEU as standard evaluation measures. The highest score reported was ROUGE-1 = 75.21. We believe that our efforts in Urdu paraphrase generation research would be fruitful not only for the Urdu language but also for similar work in other low-resource languages.
dc.identifier.urihttps://repository.cuilahore.edu.pk/handle/123456789/1751
dc.language.isoen
dc.publisherLibrary Information Services, COMSATS University Islamabad, Lahore Campus
dc.relation.ispartofseriesLHR TP 8681
dc.subjectDr. Muhammad Sharjeel
dc.subjectsp22
dc.subjectDepartment of Computer Science
dc.subjectTECHNOLOGY::Information technology::Computer science
dc.subjectUrdu Sentential
dc.subjectNatural Language Processing (NLP)
dc.subjectInformation Retrieval (IR)
dc.titleUrdu Sentential Paraphrase Generation using GAN
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
8681.pdf
Size:
405.82 KB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
319 B
Format:
Item-specific license agreed to upon submission
Description:

Collections