# Application of elementary probability models for text homogeneity and segmentation: A case study of Bible

**Authors:** Berhane Abebe, Roy Cerqueti, Roy Cerqueti, Roy Cerqueti, Roy Cerqueti

PMC · DOI: 10.1371/journal.pone.0303432 · 2024-06-07

## TL;DR

This study uses probability models to analyze the structure of Bible translations in Tigrigna, Amharic, and English, finding that they are made up of heterogeneous segments.

## Contribution

The paper applies newly developed probability models to detect text homogeneity and change points in religious texts.

## Key findings

- Bible translations in Tigrigna, Amharic, and English show heterogeneous concatenation of different books or genres.
- The Pauline letters in the English Bible are found to be heterogeneous, composed of two homogeneous segments.

## Abstract

For the purpose of this study, A statistical test of Biblical books was conducted using the recently discovered probability models for text homogeneity and text change point detection. Accordingly, translations of Biblical books of Tigrigna and Amharic (major languages spoken in Eritrea and Ethiopia) and English were studied. A Zipf-Mandelbrot distribution with a parameter range of 0.55 to 0.88 was obtained in these three Bibles. According to the statistical analysis of the texts’ homogeneity, the translation of Bible in each of these three languages was a heterogeneous concatenation of different books or genres. Furthermore, an in-depth examination of the text segmentation of prat of a single genre—the English Bible letters revealed that the Pauline letters are heterogeneous concatenations of two homogeneous segments.

## Full-text entities

- **Diseases:** CHANGE (MESH:D009402), PONE-D-23-20994HOMOGENEITY (OMIM:615816)
- **Species:** Homo sapiens (human, species) [taxon 9606]

## Figures

48 figures with captions in the complete paper: https://tomesphere.com/paper/PMC11161094/full.md

---
Source: https://tomesphere.com/paper/PMC11161094