Department of Computer Science

Permanent URI for this communityhttps://repository.cuilahore.edu.pk/handle/123456789/16

Browse

Search Results

Now showing 1 - 1 of 1
  • Item
    Cross-Lingual Abstractive Text Summarization for English-Urdu Language Pair
    (Library Information Services, COMSATS University Islamabad, Lahore Campus, 2021) Duaa-e-Nadeem; SP20-RCS-001; LHR TP 7600; Dr. Rao Muhammad Adeel Nawab
    Cross-Lingual Abstractive Text Summarization (CLATS) is a process to generate an abstractive summary in Language B (target language) of the given document in Language A (source language). CLATS has the potential to allow cross-lingual information access for millions of people around the world who do not speak or understand languages with a broad online presence and it makes documents originally written in native languages available to a wide audience. The popularity of digital documents has resulted in a rapid increase in the amount of information available online. This massive volume of data, however, cannot be manually analyzed. A summary of a big written document allows individuals to quickly grasp the topic's material without reading the entire document. CLATS plays a vital role in news summarization, literature summarization, timeline/event summarization and summarization of Wikipedia or any information source articles for better understanding. This research aims at exploring the problem of CLATS. Previously, CLATS systems have been developed for different language pair like Chinese-English, German-English, Arabic-English and vice versa but this problem is not explored for English-Urdu language pair, even though Urdu is a widely spoken South Asian language. The twofold aim of this research is to fill the gap: (1) develop a large benchmark corpus CLATS for English-Urdu language pair and (2) develop, apply, evaluate and compare state-of the-art deep learning techniques on proposed corpus for CLATS for English-Urdu language pair. For developing CLATS-EU corpus, News Summary and All the News are used as base corpus. The following steps were taken to create the proposed corpus: (1) automatic translation of base corpus (2) manual inspection and correction of testing data of base corpus (3) corpus standardization After development of CLATS-EU corpus, Deep learning sequence to sequence techniques (RNN, LSTM, GRU, BRNN, BLSTM and BGRU) are developed by using CLATS-EU corpus. Then RNN, LSTM and GRU models are implemented by applying attention mechanism. ROUGE1, ROUGE2 and ROUGEL are used as evaluation x measures to evaluate the performance of the techniques that we have developed. For CLATS, GRU gave the best results on our proposed CLATS-EU corpus