Intertextuality
Repeated sequences
Compare two selected works by exact form n-grams.
Texts and method
Select at least two works in the catalogue.
Exposed stop-list
A sequence must contain at least two tokens outside this editable list. Stop-words remain part of the exact sequence but do not increase its improbability score.
Lemma n-grams activate only from the complete corpus-wide Sôtêr snapshot. Competing predictions are preserved as sets and never resolved arbitrarily. Semantic similarity remains separately labelled and is not simulated.
Results
Shared sequences
Scoring method
The score is the sum of the negative base-10 logarithms of add-one-smoothed token probabilities in the two compared texts, excluding the visible stop-list. A larger score means that the shared sequence is less probable under this explicit independence model. It is a ranking aid, not proof of dependence or direction of borrowing.