Voynich StudiesBeinecke MS 408
RU EN IT

Language claims

Hebrew by anagram: why the algorithm named a language and the reading still did not work

In 2016 an algorithm went through hundreds of languages, picked Hebrew, and then produced the first phrase of the manuscript. We separate two different claims - identifying the language and the reading itself - and show why they hold each other up in a circle.

Verdict REJECTED 27.07.2026

Hebrew by anagram: why the algorithm named a language and the reading still did not work
Folio f77r. Beinecke Library, Yale (public domain)

This work differs from most "decipherments" in having been done by careful people and published in a peer-reviewed venue in computational linguistics. That is precisely why it is worth taking apart on the merits: the error here lies not in carelessness but in the structure of the inference itself.

Two different claims

The paper in fact contains two independent steps, and they must not be conflated.

Step one: identifying the language. The algorithm compares the frequency profile of the text with the profiles of several hundred languages and picks the closest. On the Voynich text Hebrew won.

Step two: the reading. The authors supposed that the letters within words had been rearranged into alphabetical order, restored the original order, put the result through machine translation and smoothed it a little by hand. Out came a phrase that looked meaningful - about recommendations to a priest and to the people of the house.

Each step deserves separate treatment.

Why language identification does not work here

Identifying a language by frequencies is a working instrument when the text is not enciphered. It compares the distribution of signs and finds the nearest match. But a cipher is exactly what rewrites that distribution.

We have a direct measurement of this, and a recent one. We put eight languages - Latin, Occitan, Manchu, medieval and modern Chinese, classical Japanese, sixteenth-century Nahuatl - through the Voynich cipher model at an equal volume of text. The result: all eight leave practically the same fingerprint, and the difference between languages disappears. The details are in our review of the East Asian version, where we also retracted an earlier conclusion of our own.

Hence the consequence: under a cipher of this type the original language does not show through. Any ranking of languages built on an enciphered text ranks artefacts of the cipher rather than languages. Hebrew winning such a ranking is not evidence about Hebrew.

There is a second and more general measurement of ours as well: the attempt to recover the key from the internal statistics failed so thoroughly that the real text came out worse than random noise. The trace of the original letters has been erased. A language cannot be identified from an erased trace.

Why an anagram reading cannot be tested

The second step suffers from the same disease as most "readings": too much freedom.

If letters may be rearranged within a word, then several meaningful words can nearly always be assembled from a set of signs. The choice between them is made by a human being, according to what fits the phrase better. Then machine translation comes in, which is itself inclined to produce connected text even out of rubbish, followed by a final tidying by hand.

At each of the three stages the system gains a degree of freedom. Taken together, they will read almost any set of signs. And so a successfully read phrase is not evidence: a hypothesis that forbids no outcomes cannot be tested.

There is an internal problem too: the anagrams are restored from the ciphertext, that is, from the very material whose structure we are trying to explain. This is a circle - the assumption of rearrangement is justified by the fact that rearrangement yields a meaningful text, while the meaningfulness is supplied by the freedom of choice that the assumption itself grants.

What a testable version would look like

For comparison, here is what would have to be presented:

  • a table of correspondences fixed before translation and not varying from word to word;
  • a reading of a stretch of text that took no part in the tuning;
  • a regular grammar: the same endings in the same positions;
  • an independent prediction - for instance, that a particular word will occur on the page with a

particular drawing.

Not one of these points has been met in any of the known "decipherments" - not in the Hebrew one, not in the proto-Romance one, not in the Aztec one.

What we claim and what we do not

We are not saying "Hebrew is ruled out". Our statement is stricter and duller: the method as presented gives no grounds either for or against, because it measures the wrong thing.

The original language cannot be named today by any available means, and that applies to Hebrew exactly as it applies to Latin, to Occitan or to Nahuatl. Such is the price of a cipher that erases the trace of the letters.

No decipherment of the Voynich manuscript exists. A coincidence is not a reading.

FAQ

What did Hauer and Kondrak do?

They built an algorithm that compares the frequency profile of a text with the profiles of several hundred languages. On the Voynich text the algorithm picked Hebrew. The authors then supposed that the letters within words had been rearranged, restored the order, and obtained a phrase that looked meaningful.

Surely identifying the language by algorithm is a strong argument?

Only if the text is not enciphered. The algorithm compares the distribution of signs, and a cipher rewrites exactly that distribution. Our measurements show that under a cipher of this type the original language does not show through at all, so a win by any language in such a ranking means nothing.

What is wrong with the anagrams?

The freedom of rearrangement is too great. Out of a set of letters one can nearly always assemble several meaningful words, and the choice is made by what seems to fit. There is nothing with which to test a reading of that kind: it forbids no outcomes.

Does that mean Hebrew is ruled out entirely?

No. Our statement is stricter than that: this method gives no grounds either for or against. The original language cannot be named today by any available means - and that applies to Hebrew exactly as it applies to Latin and to Nahuatl.

← All hypotheses