Developing a Corpus of Paraphrased Plagiarized Documents for Urdu Language

dc.contributor.authorIqra Sharif
dc.contributor.authorBCS-SP11-42
dc.contributor.authorDr. RAO MUHAMMAD ADEEL NAWAB
dc.contributor.authorLHR TP 6919
dc.date.accessioned2026-02-25T06:01:05Z
dc.date.issued2015-11-20
dc.description.abstractPlagiarism (unacknowledged reuse of text), is a wide spread problem around the world, particularly in academia. Research has shown that it is easy to detect exact copy however; it is hard to detect plagiarism when the original text has been heavily paraphrased. The focus of this research is on paraphrased plagiarism detection. This FYP report describes our work to investigate the problem of paraphrase plagiarism in Urdu language. The main focus of this project was exploring the problem of extrinsic plagiarism detection specially when the rewritten text has been paraphrased. Our first aim is develop Urdu corpus with instances of paraphrased plagiarised documents and non-plagiarised documents. Once the corpus was built, Natural Language Processing (NLP) algorithms were applied to detect the similarity between the original document and the suspicious document. The problem of discriminating paraphrased plagiarised documents from non-plagiarised documents is casted as a supervised classification task using existing Machine Learning (ML) techniques. Two popular and well know techniques were explored to identify paraphrased plagiarism: (1) N-gram overlap and (2) Longest Common Subsequence (LCS). The similarity between plagiarised and non-plagiarised documents was computed using overlap similarity score for unigrams, bigrams, trigrams, four-grams and five-grams and LCS. These similarity scores were used as features for classification. Evaluation was carried out using precision, recall and F1 measure. Results showed that it is not trivial to detect plagiarism created by paraphrasing the original text.
dc.identifier.urihttps://repository.cuilahore.edu.pk/handle/123456789/2219
dc.publisherLibrary Information Services, COMSATS University Islamabad, Lahore Campus.
dc.relation.ispartofseriesLHR TP 6919; LHR TP 6919
dc.subjectDeveloping a Corpus of Paraphrased Plagiarized Documents for Urdu Language
dc.subjectTECHNOLOGY::Information technology::Computer science::Computer science
dc.subjectSP11
dc.titleDeveloping a Corpus of Paraphrased Plagiarized Documents for Urdu Language
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
Developing a Corpus of Paraphrased Plagiarized Documents for Urdu Language (report).docx
Size:
219.12 KB
Format:
Microsoft Word XML

License bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
319 B
Format:
Item-specific license agreed to upon submission
Description: