conference-paper

The Use of Stochastic Models in the Analysis of Vast English Literary Data Corpora

Research footprint

At a glance

Citations
0
References
26
Comments
0
Paper overview

Öz

In the stylometric analysis of literary works, there are primarily specific aims, for example, at resolving the authorship, the period of the work, style, motif and purpose. These tasks, in fact, can be placed in a machine learning framework, where the quantitative features of known passages are initially learned and analyzed, and then compared with the corresponding metrics in disputable or doubtful passages. While in the past, relative laborious manual approaches relying on the human judgment of selected literary experts are mainly employed, they have shown to error-prone since human experts are unable to systematically process and methodically integrate the varied features of vast volumes of literary data. By combining a big data approach with the application of sophisticated stochastic models, more efficient, dependable and defensible decisions can be made. In this paper, we introduce two models for the quantitative analysis of the characteristics of literary work. The first is a Bernoulli model, which assumes the underlying quantitative features that are independent of each other, and the second is a Markov model, which is particularly appropriate for the analysis of poetic styles. We derive probabilistic measures of passage characteristics, and provide their illustrations in literary works in the English language. Experiments are performed the results of which corroborate the approach.

Record transparency

Publication details

DOI
10.1109/bigdia51454.2020.00052
OpenAlex
W3144158305
Document type
conference-paper
Language
EN
Last metadata update
Community

Comments

Oturum Açın to join the discussion.

  1. No comments yet. Start the discussion.