Cross-Lingual Abstractive Text Summarization for English-Urdu Language Pair

dc.contributor.authorDuaa-e-Nadeem
dc.contributor.authorSP20-RCS-001
dc.contributor.authorLHR TP 7600
dc.contributor.authorDr. Rao Muhammad Adeel Nawab
dc.date.accessioned2026-02-16T04:23:35Z
dc.date.issued2021
dc.description.abstractCross-Lingual Abstractive Text Summarization (CLATS) is a process to generate an abstractive summary in Language B (target language) of the given document in Language A (source language). CLATS has the potential to allow cross-lingual information access for millions of people around the world who do not speak or understand languages with a broad online presence and it makes documents originally written in native languages available to a wide audience. The popularity of digital documents has resulted in a rapid increase in the amount of information available online. This massive volume of data, however, cannot be manually analyzed. A summary of a big written document allows individuals to quickly grasp the topic's material without reading the entire document. CLATS plays a vital role in news summarization, literature summarization, timeline/event summarization and summarization of Wikipedia or any information source articles for better understanding. This research aims at exploring the problem of CLATS. Previously, CLATS systems have been developed for different language pair like Chinese-English, German-English, Arabic-English and vice versa but this problem is not explored for English-Urdu language pair, even though Urdu is a widely spoken South Asian language. The twofold aim of this research is to fill the gap: (1) develop a large benchmark corpus CLATS for English-Urdu language pair and (2) develop, apply, evaluate and compare state-of the-art deep learning techniques on proposed corpus for CLATS for English-Urdu language pair. For developing CLATS-EU corpus, News Summary and All the News are used as base corpus. The following steps were taken to create the proposed corpus: (1) automatic translation of base corpus (2) manual inspection and correction of testing data of base corpus (3) corpus standardization After development of CLATS-EU corpus, Deep learning sequence to sequence techniques (RNN, LSTM, GRU, BRNN, BLSTM and BGRU) are developed by using CLATS-EU corpus. Then RNN, LSTM and GRU models are implemented by applying attention mechanism. ROUGE1, ROUGE2 and ROUGEL are used as evaluation x measures to evaluate the performance of the techniques that we have developed. For CLATS, GRU gave the best results on our proposed CLATS-EU corpus
dc.identifier.urihttps://repository.cuilahore.edu.pk/handle/123456789/1641
dc.language.isoen
dc.publisherLibrary Information Services, COMSATS University Islamabad, Lahore Campus
dc.relation.ispartofseriesLHR TP 7600
dc.subjectDr. Rao Muhammad Adeel Nawab
dc.subjectSp20
dc.subjectDepartment of Computer Science
dc.subjectComputer Science
dc.subjectMATHEMATICS::Applied mathematics::Theoretical computer science
dc.subjectText Summarization
dc.titleCross-Lingual Abstractive Text Summarization for English-Urdu Language Pair
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
SP20-RCS-001_Duaa-e-Nadeem_Thesis.pdf
Size:
3.23 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
319 B
Format:
Item-specific license agreed to upon submission
Description:

Collections