Development of the linguometric method for automatic identification of the author of text content based on statistical analysis of language diversity coefficients

Abstract

<p>We have developed the linguometric method for algorithmic support of content monitoring processes to solve the problem of the automatic identification of the author of the Ukrainian text content based on the technology of statistical analysis of the language diversity coefficients. The decomposition of the method for identification of the author based on the analysis of such speech factors as lexical diversity, degree (measure) of syntactic complexity, speech coherence, indexes of exclusivity and concentration of a text was performed. Such parameters of the author’s style as the number of words in the specified text, the total number of words in this text, the number of sentences, the number of prepositions, the number of conjunctions, the number of words with the frequency of 1, the number of words with the frequency of 10 and more were analyzed. The features of the developed methods are the adaptation of the morphological and syntactic analysis of lexical units to the peculiarities of the structures of Ukrainian words/texts. That is, when analyzing linguistic units of the word type, their belonging to a part of speech and declension within this part of speech was taken into account. For this, the flections of these words for their classification, separation of the base for the formation of the corresponding alphabetic-frequency dictionaries were analyzed. Filling these dictionaries was subsequently taken into consideration at the following stages of the identification of the authorship of a text, such as the calculation of parameters and coefficients of the author's speech. Syntactic words (stop or anchor) words are most essential for an individual style of an author, as they are not related to the subject and content of the publication. We compared the results in a set of 200 one-author papers in the technical area of more than 100 different authors over the period of 2001–2017 to determine if and how the coefficients of diversity of a text of these authors change within different periods of time. It was found that for the selected experimental base of more than 200 papers, the best results according to the density criterion are reached by the method for analysis of an article without the initial compulsory information, such as abstracts and keywords in different languages, as well as the list of literature.</p>

Authors and Affiliations

Vasyl Lytvyn, Victoria Vysotska, Petro Pukach, Zinovii Nytrebych, Ihor Demkiv, Roman Kovalchuk, Nadiia Huzyk

Keywords

Related Articles

Design and study of equipment for accepting and drying soya seeds with high moisture content

<p>The studied designs of existing equipment for post-harvest processing do not ensure careful reception and drying of high-moisture soya seeds which leads to a decrease in quality and yield loss. To solve this problem,...

Advancement of a long arithmetic technology in the construction of algorithms for studying linear systems

<p>We have advanced the application of algorithms within a method of basic matrices, which are equipped with the technology of long arithmetic to improve the precision of performing the basic operations in the course of...

Ray tracing synthesis of images of triangulated surfaces smoothed by the spherical interpolation method

<span lang="EN-US">The problem of imaging by ray tracing of triangulated surfaces smoothed by the spherical interpolation method was solved. The method of spherical interpolation was mainly designed to interpolate the tr...

Evaluation of dynamic properties of gas pumping units according to the results of experimental researches

Experimental studies of the dynamic properties of gas pumping units (GPU) of various types, which allowed us to obtain GPU acceleration curves and determine the parameters of the transfer function through various transmi...

Modeling a thermal conductivity process under the action of flame on the wall of fire­retardant reed

<p>Creating environmentally friendly flame-retardant materials for natural inflammable roof structures will make it possible to control the processes of thermal stability and physical-chemical properties of a protective...

Download PDF file
  • EP ID EP528152
  • DOI 10.15587/1729-4061.2018.142451
  • Views 82
  • Downloads 0

How To Cite

Vasyl Lytvyn, Victoria Vysotska, Petro Pukach, Zinovii Nytrebych, Ihor Demkiv, Roman Kovalchuk, Nadiia Huzyk (2018). Development of the linguometric method for automatic identification of the author of text content based on statistical analysis of language diversity coefficients. Восточно-Европейский журнал передовых технологий, 5(2), 16-28. https://europub.co.uk/articles/-A-528152