Development of the linguometric method for automatic identification of the author of text content based on statistical analysis of language diversity coefficients

Abstract

<p>We have developed the linguometric method for algorithmic support of content monitoring processes to solve the problem of the automatic identification of the author of the Ukrainian text content based on the technology of statistical analysis of the language diversity coefficients. The decomposition of the method for identification of the author based on the analysis of such speech factors as lexical diversity, degree (measure) of syntactic complexity, speech coherence, indexes of exclusivity and concentration of a text was performed. Such parameters of the author’s style as the number of words in the specified text, the total number of words in this text, the number of sentences, the number of prepositions, the number of conjunctions, the number of words with the frequency of 1, the number of words with the frequency of 10 and more were analyzed. The features of the developed methods are the adaptation of the morphological and syntactic analysis of lexical units to the peculiarities of the structures of Ukrainian words/texts. That is, when analyzing linguistic units of the word type, their belonging to a part of speech and declension within this part of speech was taken into account. For this, the flections of these words for their classification, separation of the base for the formation of the corresponding alphabetic-frequency dictionaries were analyzed. Filling these dictionaries was subsequently taken into consideration at the following stages of the identification of the authorship of a text, such as the calculation of parameters and coefficients of the author's speech. Syntactic words (stop or anchor) words are most essential for an individual style of an author, as they are not related to the subject and content of the publication. We compared the results in a set of 200 one-author papers in the technical area of more than 100 different authors over the period of 2001–2017 to determine if and how the coefficients of diversity of a text of these authors change within different periods of time. It was found that for the selected experimental base of more than 200 papers, the best results according to the density criterion are reached by the method for analysis of an article without the initial compulsory information, such as abstracts and keywords in different languages, as well as the list of literature.</p>

Authors and Affiliations

Vasyl Lytvyn, Victoria Vysotska, Petro Pukach, Zinovii Nytrebych, Ihor Demkiv, Roman Kovalchuk, Nadiia Huzyk

Keywords

Related Articles

Development of cleaning methods complex of industrial gas pipelines based on the analysis of their hydraulic efficiency

<p>The majority of gas and gas condensate fields of Ukraine are developed by pressure depletion, which makes it possible to stabilize production only in conditions of low working pressures at the wellhead. In turn, the w...

A study of initial stages for formation of carbon condensates on copper

<p>In the CVD method, samples of carbon condensates were obtained under special conditions (low substrate temperature and short growth times). The use of special technological conditions makes it possible to study the in...

Development of methods for the analysis of functional requirements to an information system for consistency and illogicality

<p>The order of work aimed at analyzing the requirements is considered, taking into account application of the service approach to development of IS, models and methods of development of representations of functional req...

Influence of the CaO-containing modifiers on the properties of alkaline alyumosilicate binders

<p>The basis for ensuring the resistance of artificial stone based on alkaline aluminosilicate binders to variable environmental conditions is the formation of zeolite- and mica-like hydrate neo-formations.</p><p>It is p...

Comparison of products of whey proteins concentrate proteolysis, obtained by different proteolytic preparations

<p>An important source of bioactive peptides is hydrolyzed products based on milk whey: hypoallergenic products, hydrolyzates for baby food, and products for athletes. However, in their production, proteolytic preparatio...

Download PDF file
  • EP ID EP528152
  • DOI 10.15587/1729-4061.2018.142451
  • Views 80
  • Downloads 0

How To Cite

Vasyl Lytvyn, Victoria Vysotska, Petro Pukach, Zinovii Nytrebych, Ihor Demkiv, Roman Kovalchuk, Nadiia Huzyk (2018). Development of the linguometric method for automatic identification of the author of text content based on statistical analysis of language diversity coefficients. Восточно-Европейский журнал передовых технологий, 5(2), 16-28. https://europub.co.uk/articles/-A-528152