Text Reuse Detection for the Urdu Language at Lexical Level

dc.contributor.authorAyesha Noreen
dc.contributor.authorFA21-RCS-026
dc.contributor.authorLHR TP 8461
dc.contributor.authorDr. Rao Muhammad Adeel Nawab
dc.date.accessioned2026-02-16T09:19:47Z
dc.date.issued2023
dc.description.abstractText reuse is a process of creating new text(s) from pre-existing text(s). In recent years, Urdu Text Reuse Detection (UTRD) has gained the attention of researchers because the text is readily available in digital format all over the internet and can be copied or paraphrased from another source without proper attribution, which makes it easier to reuse but hard to detect. In previous studies, the problem of UTRD has been explored at the sentence level, sentence/passage level, and document level, along with benchmark corpora and methods. However, the problem of UTRD has not been explored at the Lexical level with respect to corpora and methods. To fulfill this research gap, this research study has developed a large benchmark manually annotated corpus of 22,184 text pairs at two levels of rewrite: 1) Derived = 8,660 and 2) Non-Derived = 13,524. Secondly, this research study has developed, applied, evaluated, and compared baseline approaches (Uni-gram Overlap and Word Embedding-based approaches) with proposed transformer-based approaches on the proposed UTRD-Lex-22 corpus. The proposed transformers-based approach (using a combination of 7 different Sentence Transformers including: distiluse-base-multilingual-cased-v1, LaBSE, paraphrase-multilingual-MiniLM-L12-v2, paraphrase-multilingual-mpnet-base-v2, xlm-r-100langs-bert-base-nli-mean-tokens, xlm-r-bert- base-nli-stsb-mean-tokens, xlm-r-distilroberta-base-paraphrase-v1) outperforms with an F1 score of 0.69 compared to the best results obtained using feature fusion approach (F1 = 0.70).
dc.identifier.urihttps://repository.cuilahore.edu.pk/handle/123456789/1707
dc.language.isoen
dc.publisherLibrary Information Services, COMSATS University Islamabad, Lahore Campus
dc.relation.ispartofseriesLHR TP 8461
dc.subjectFa21
dc.subjectDr. Rao Muhammad Adeel Nawab
dc.subjectDepartment of Computer Science
dc.subjectComputer Science
dc.subjectUrdu Text Reuse Detection (UTRD)
dc.titleText Reuse Detection for the Urdu Language at Lexical Level
dc.typeThesis

Files

Original bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
Final-Thesis - Text reuse detection for the urdu language at lexical level - FA21-RCS-026.docx
Size:
985.8 KB
Format:
Microsoft Word XML

License bundle

Now showing 1 - 1 of 1
No Thumbnail Available
Name:
license.txt
Size:
319 B
Format:
Item-specific license agreed to upon submission
Description:

Collections