Developing a Corpus of Sentential Paraphrases for Urdu Language

No Thumbnail Available

Date

2015-11-20

Journal Title

Journal ISSN

Volume Title

Publisher

Library Information Services, COMSATS University Islamabad, Lahore Campus.

Abstract

The main aim of this project was to explore the problem of paraphrasing at sentence level. Our first task was to develop an Urdu corpus with instances of paraphrase and non-paraphrase sentence pairs. The corpus created contains examples of sentence level paraphrase news pairs from journalism. After developing corpus we applied the existing Natural Language Processing (NLP) algorithms to detect the similarity between the original and rewritten (paraphrased) text. The problem of discriminating paraphrase pairs from non-paraphrase ones casted as a supervised classification task using existing Machine Learning (ML) algorithms. We evaluated the effect of paraphrase and non-paraphrase pairs in the dataset. N-gram overlap and Longest Common Subsequence (LCS) two popular and effective approaches for paraphrase detection were applied on the dataset. Containment similarity score for unigrams, bigrams, trigrams, four-grams and five-grams along with LCS were used as features for classification. Evaluation was carried out using precision, recall and F1 measure. Results showed that it is not trivial to detect paraphrasing created using manual methods.

Description

Keywords

Developing a Corpus of Sentential Paraphrases for Urdu Language, Computer science, SP11, problem of paraphrasing, Urdu corpus

Citation

Endorsement

Review

Supplemented By

Referenced By