Developing a Corpus of Sentential Paraphrases for Urdu Language
No Thumbnail Available
Date
2015-11-20
Journal Title
Journal ISSN
Volume Title
Publisher
Library Information Services, COMSATS University Islamabad, Lahore Campus.
Abstract
The main aim of this project was to explore the problem of paraphrasing at sentence level. Our first task was to develop an Urdu corpus with instances of paraphrase and non-paraphrase sentence pairs. The corpus created contains examples of sentence level paraphrase news pairs from journalism. After developing corpus we applied the existing Natural Language Processing (NLP) algorithms to detect the similarity between the original and rewritten (paraphrased) text. The problem of discriminating paraphrase pairs from non-paraphrase ones casted as a supervised classification task using existing Machine Learning (ML) algorithms.
We evaluated the effect of paraphrase and non-paraphrase pairs in the dataset. N-gram overlap and Longest Common Subsequence (LCS) two popular and effective approaches for paraphrase detection were applied on the dataset. Containment similarity score for unigrams, bigrams, trigrams, four-grams and five-grams along with LCS were used as features for classification. Evaluation was carried out using precision, recall and F1 measure. Results showed that it is not trivial to detect paraphrasing created using manual methods.
Description
Keywords
Developing a Corpus of Sentential Paraphrases for Urdu Language, Computer science, SP11, problem of paraphrasing, Urdu corpus