• DocumentCode
    2338672
  • Title

    Algorithm of the longest commonly consecutive word for Plagiarism detection in text based document

  • Author

    Sediyono, Agung ; Ku-Mahamud, Ku Ruhana

  • Author_Institution
    Inf. Eng. Dept., Universitas Trisakti, Jakarta
  • fYear
    2008
  • fDate
    13-16 Nov. 2008
  • Firstpage
    253
  • Lastpage
    259
  • Abstract
    Plagiarism is a form of academic misconduct which has increased with the easy access to obtain information through electronic documents and the Internet. The problem of finding document plagiarism in full text document can be viewed as a problem of finding the longest common parts of strings. Moreover, the detection system has to be capable to determine and visualize not only the common parts but also the location of the common parts in both the source and the observed document. Unlike previous research, this paper proposes a numerical based comparison algorithm that is comparable in the computation time without loosing the word order of common parts. Based on the experiment, the proposed algorithm outperforms the suffix tree in the length of observed paragraph below one hundred words.
  • Keywords
    Internet; document handling; trees (mathematics); Internet; academic misconduct; consecutive word; document plagiarism; electronic documents; plagiarism detection; suffix tree; text based document; Art; Design for experiments; Educational institutions; Electronic mail; Face detection; Filters; Informatics; Internet; Plagiarism; Visualization;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Digital Information Management, 2008. ICDIM 2008. Third International Conference on
  • Conference_Location
    London
  • Print_ISBN
    978-1-4244-2916-5
  • Electronic_ISBN
    978-1-4244-2917-2
  • Type

    conf

  • DOI
    10.1109/ICDIM.2008.4746827
  • Filename
    4746827