Muhammad Umar SarwarSp11-Bcs-148Dr Muhammad Rao Adeel NawabLHR TP 69162026-02-252015-11-20https://repository.cuilahore.edu.pk/handle/123456789/2270The main aim of this project was to explore the problem of paraphrasing at sentence level. Our first task was to develop an Urdu corpus with instances of paraphrase and non-paraphrase sentence pairs. The corpus created contains examples of sentence level paraphrase news pairs from journalism. After developing corpus we applied the existing Natural Language Processing (NLP) algorithms to detect the similarity between the original and rewritten (paraphrased) text. The problem of discriminating paraphrase pairs from non-paraphrase ones casted as a supervised classification task using existing Machine Learning (ML) algorithms. We evaluated the effect of paraphrase and non-paraphrase pairs in the dataset. N-gram overlap and Longest Common Subsequence (LCS) two popular and effective approaches for paraphrase detection were applied on the dataset. Containment similarity score for unigrams, bigrams, trigrams, four-grams and five-grams along with LCS were used as features for classification. Evaluation was carried out using precision, recall and F1 measure. Results showed that it is not trivial to detect paraphrasing created using manual methods.Developing a Corpus of Sentential Paraphrases for Urdu LanguageComputer scienceSP11problem of paraphrasingUrdu corpusDeveloping a Corpus of Sentential Paraphrases for Urdu LanguageThesis