Development of the linguometric method for automatic identification of the author of text content based on statistical analysis of language diversity coefficients

Abstract

<p>We have developed the linguometric method for algorithmic support of content monitoring processes to solve the problem of the automatic identification of the author of the Ukrainian text content based on the technology of statistical analysis of the language diversity coefficients. The decomposition of the method for identification of the author based on the analysis of such speech factors as lexical diversity, degree (measure) of syntactic complexity, speech coherence, indexes of exclusivity and concentration of a text was performed. Such parameters of the author’s style as the number of words in the specified text, the total number of words in this text, the number of sentences, the number of prepositions, the number of conjunctions, the number of words with the frequency of 1, the number of words with the frequency of 10 and more were analyzed. The features of the developed methods are the adaptation of the morphological and syntactic analysis of lexical units to the peculiarities of the structures of Ukrainian words/texts. That is, when analyzing linguistic units of the word type, their belonging to a part of speech and declension within this part of speech was taken into account. For this, the flections of these words for their classification, separation of the base for the formation of the corresponding alphabetic-frequency dictionaries were analyzed. Filling these dictionaries was subsequently taken into consideration at the following stages of the identification of the authorship of a text, such as the calculation of parameters and coefficients of the author's speech. Syntactic words (stop or anchor) words are most essential for an individual style of an author, as they are not related to the subject and content of the publication. We compared the results in a set of 200 one-author papers in the technical area of more than 100 different authors over the period of 2001–2017 to determine if and how the coefficients of diversity of a text of these authors change within different periods of time. It was found that for the selected experimental base of more than 200 papers, the best results according to the density criterion are reached by the method for analysis of an article without the initial compulsory information, such as abstracts and keywords in different languages, as well as the list of literature.</p>

Authors and Affiliations

Vasyl Lytvyn, Victoria Vysotska, Petro Pukach, Zinovii Nytrebych, Ihor Demkiv, Roman Kovalchuk, Nadiia Huzyk

Keywords

Related Articles

Development of signal converter of thermal sensors based on combination of thermal and capacity research methods

<p class="a"><span lang="EN-US">The problem of functional integration of thermal and capacity research methods, which provides the possibility of realizing a new generation of analog front­end of the Internet of Things i...

Improvement of safety of autonomous electrical installations by implementing a method for calculating the electrolytic grounding electrodes parameters

We have solved the task of safety improvement in the grounding process of autonomous mobile electrical installations. Existing procedures for the calculation of normalized resistance of grounding electrodes in electric i...

Effect of thermal field distribution in the layered structure of a heating floor on the temperature of its surface

<p>We propose a method for creating optimum temperature microclimate modes at livestock facilities of different functional purpose by using a multi-layer heating floor. A structural mathematical model was constructed tha...

Studying the influence of hydrothermal treatment parameters on the properties of wheat flour in the technology of a croquette mass

<p class="a">The paper reports a study into the influence of parameters (temperature and duration) of sautéing wheat flour on its water­absorbing capacity at hydrothermal treatment. It was established that heating the sa...

Energy absorbers on the steel plate – rubber laminate after deformable projectile impact

<p>The ability of energy absorption can be used to measure the strength of material against ballistic impact. This paper aims to analyze the rubber plated energy absorption plate that was shot with deformable projectiles...

Download PDF file
  • EP ID EP528152
  • DOI 10.15587/1729-4061.2018.142451
  • Views 43
  • Downloads 0

How To Cite

Vasyl Lytvyn, Victoria Vysotska, Petro Pukach, Zinovii Nytrebych, Ihor Demkiv, Roman Kovalchuk, Nadiia Huzyk (2018). Development of the linguometric method for automatic identification of the author of text content based on statistical analysis of language diversity coefficients. Восточно-Европейский журнал передовых технологий, 5(2), 16-28. https://europub.co.uk/articles/-A-528152