Voynich StudiesBeinecke MS 408
RU EN IT

Hoax and gibberish

Voynich entropy: the argument most often misused

"The entropy is too low, so this cannot be gibberish" is the commonest argument in popular retellings, and it is wrong. We set out what has actually been measured, which numbers are real, and what follows from them.

Verdict BACKGROUND 27.07.2026

Voynich entropy: the argument most often misused
Folio f108r. Beinecke Library, Yale (public domain)

Everyone who writes about the manuscript at all writes about its entropy - and almost always with the same logical error. Let us go through the numbers that stand behind the word "entropy", which of them are real, and what conclusion may be drawn from them.

Two different quantities that are constantly confused

Sign entropy (written h1) answers the question: how varied are the letters themselves? If the alphabet has thirty signs and all of them occur about equally often, the figure is high; if a single sign takes up half the text, it is low.

Conditional entropy (h2) answers a different question: how easy is it to guess the next sign when the previous one is known? In English, q is almost always followed by u, and that lowers the figure. The tighter the constraints on combination, the lower the number.

These are different things, and the confusion between them is the source of most of the false claims in popular retellings.

The real numbers

Here is what has been measured on the full transliteration of the manuscript:

QuantityVoynichNatural languages
Entropy of a single glyph (h1)3.86 bitscomparable
Conditional entropy (h2)about 2.1-2.2 bits2.96-3.23

It is the second row that matters. What is anomalous is not the alphabet but the combinations. Voynich signs follow one another far more predictably than the letters of any natural language tested.

While we are here: the figures "4.0-4.3 against 4.5-4.8" that circulate in reviews match neither our measurements nor the known publications - they are exactly the result of mixing h1 with h2.

What low conditional entropy does mean

It means a rigid internal structure, and nothing more than that.

In our model such a structure is created by a syllabic cipher: one unit of plaintext is written with two or three signs from a restricted set, so that combinations recur and predictability rises. We tested this by direct construction: a Latin text with a conditional entropy of 3.22, passed through the model, falls to roughly 2.1 - that is, precisely into the observed region.

The same thing explains the second peculiarity, the short and uniform words: they are set by the cipher, not by the original language.

What low entropy does NOT mean

Here is where the main error of popular accounts lies: "the entropy is too low, therefore the text is meaningful rather than a meaningless set of signs".

This is wrong. Mechanical generators produce the same low entropy.

  • A grille with a table of syllables - the device with which Gordon Rugg reproduced a text resembling

the Voynich - yields predictability automatically, because the blocks repeat.

  • Copying a word already written with a small alteration - the mechanism proposed by Torsten Timm -

also yields low entropy, and on top of that Zipf's law, families of similar words and clustering of repeats.

In other words, that one number cannot distinguish a meaningful source from a mechanical one. The gibberish version has to be rejected by discriminative features - the ones a generator overshoots or undershoots. We have done that separately: the Cardan grille has been rejected in two ways, and autocopying on five measures out of six on held-out data.

The lesson for reading any review

A number on its own settles nothing. What settles matters is whether it discriminates between the competing explanations.

We stepped on an adjacent rake ourselves and admitted it publicly: our measure of "twin words" turned out to depend on the size of the corpus, and we retracted the conclusion about East Asian languages that had been built on it. The story is set out here.

The rule that follows is a simple one: before accepting a number as an argument, ask which alternative explanation it excludes. If it excludes none, it is not an argument but an ornament.

No decipherment of the Voynich manuscript exists. A coincidence is not a reading - and neither is low entropy.

FAQ

What is the entropy of the Voynich text?

About 3.86 bits per glyph, which is an ordinary figure. What is anomalous is the conditional entropy, that is, the predictability of the next sign given the previous one: about 2.1-2.2 bits against 2.96-3.23 for natural languages.

What does low conditional entropy mean?

That the text is highly predictable: knowing one sign makes the next easier to guess than in an ordinary language. It points to a rigid internal structure - in our model that structure is created by a syllabic cipher with a restricted set of combinations.

Is it true that low entropy proves the text is meaningful?

No. Mechanical generators produce the same low entropy: both a grille with a table of syllables and the copying of what has already been written. The gibberish version has to be rejected by other features, not by this number.

Why do reviews quote figures of 4.0-4.3 and 4.5-4.8?

They come from a confusion between two different quantities - the entropy of a single sign and the conditional entropy. Neither of these figures matches our measurements, and neither occurs in the known publications.

← All hypotheses