Development of Large Scale English-Urdu Machine Translation Corpus for Statistical and Neural Machine Translation Systems

dc.contributor.authorMoodser Hussain
dc.contributor.authorFA17-RCS-005
dc.contributor.authorLHR TP 5623
dc.contributor.authorDr. Rao Muhammad Adeel Nawab
dc.date.accessioned2026-02-11T08:11:34Z
dc.date.issued2019
dc.description.abstractThe process of automatically converting the text of one natural language into an other is known as Machine Translation (MT). The two state-of-the-art machine translation techniques are Statistical Machine Translation (SMT) and Neural Ma chine Translation (NMT). In both MT techniques, a large-scale parallel aligned corpus of source & target language text is the most basic requirement to develop quality models. For the Urdu language, there are minimal parallel resources de veloped until now, and these existing resources are not enough to produce some quality translation systems. This research study is majorly focusing on: (1) the development of a large-scale semi-automatically aligned English-Urdu parallel corpus of more than 150k sen tences by collecting the data from different domains. (2) the compilation of large sized monolingual data for English and Urdu language by merging existing cor pora. (3) the development of a character-separated parallel corpus for supervised transliteration (by using Wikipedia titles) (4) the evaluation of impact generated by language model and monolingual corpus on translation quality (by using differ ent data and word-order in language models training). (5) development of Phrase based SMT (PB-SMT) and Recurrent Neural Network-based NMT (RNN-NMT) systems to evaluate and compare their performance on the proposed parallel cor pus. (6) the usage of supervised transliteration and ‘minimum error rate training’ to enhance the quality of SMT. To evaluate the performance of MT systems, ‘Bilingual Evaluation Understudy (BLEU)’ is used as an evaluation measure. The best-achieved scores for English to-Urdu translations are 37.39 and 32.70 by using PBSMT and RNN-NMT, re spectively. While for Urdu-to-English translations, PBSMT and RNN-NMT yield the best score of 28.20 and 30.63 respectively
dc.identifier.urihttps://repository.cuilahore.edu.pk/handle/123456789/1456
dc.language.isoen
dc.publisherLibrary Information Services, COMSATS University Islamabad, Lahore Campus
dc.relation.ispartofseriesLHR TP 5623
dc.subjectDr. Rao Muhammad Adeel Nawab
dc.subjectfa17
dc.subjectDepartment of Computer Science
dc.subjectComputer Science
dc.subjectMachine Translation (MT)
dc.subjectStatistical Machine Translation (SMT)
dc.titleDevelopment of Large Scale English-Urdu Machine Translation Corpus for Statistical and Neural Machine Translation Systems
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
final_Thesis_FA17-RCS-005_Moodser Hussain.pdf
Size:
1.65 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
319 B
Format:
Item-specific license agreed to upon submission
Description:

Collections