Language claims
The East Asian version - and our own mistake, which we had to retract
The short words of the Voynich look like the syllabic languages of Asia. We built a structural filter, used it to rule out Chinese and Japanese - and then found that the filter was ranking file sizes, not languages.

The short, uniform words of the Voynich have long been explained by a language rather than by a cipher. The logic is simple: if the words are short, perhaps behind them stands a language in which short words are natural. Hence the two best-known versions - the linguist Jacques Guy suggested a monosyllabic Eastern language such as Chinese or Vietnamese, and Zbigniew Banasik proposed proto-Manchu in 2003.
We built an instrument so that versions like these could be tested rather than debated. And the honest way to begin this article is with how it ended: the instrument turned out to be faulty, and we retracted our own earlier conclusion about the Eastern languages ourselves.
How the structural filter works
We reconstructed the mechanism of the Voynich cipher: a syllabic homophonic cipher in which the same unit of plaintext could be written in several interchangeable ways, plus a few scribal habits. The mechanism can be run in the other direction: take a real text in a real language, pass it through the cipher model, and see whether the result matches the Voynich fingerprint on six numbers at once.
The point of the filter is not to "read" anything but to eliminate: if a language under our cipher leaves a completely different trace, the version weakens. One of the six measures was treated as the main one - the share of twin words, that is, pairs of different words standing in identical surroundings. In the Voynich such pairs are rare, 1.2%.
An important rule we kept to: corpora of the right period. For a hypothesis about a manuscript of the 1400s, modern Japanese will not do - its loanwords and its orthography belong to another era. We went specifically for Middle Chinese in Baxter's reconstruction, classical Japanese (bungo), and Manchu in Moellendorff transliteration.
What we got at first
In July this method produced a neat picture: the monosyllabic languages gave an avalanche of false twins - Middle Chinese 45%, Mandarin 18.6%, classical Japanese 17-27%, against a target of 1.2%. Of the Eastern languages only Manchu survived, at 1.2%. The conclusion read: both main East Asian branches fail structurally, and the version narrows to Manchu alone.
At the time we also noted a curious detail: "modern Japanese passes" turned out to be an artefact of modernity, and with the correct period the Japanese branch fell too. It looked like a lesson about discipline in choosing sources.
The lesson was a different one.
Where the mistake was
A few days later we were running Nahuatl through the same filter, to test the Aztec hypothesis. Nahuatl gave 32% twins. Before writing the conclusion down, we checked that measure across all the corpora in the project - and saw something that should not have been there.
| Corpus | Size | Share of twin words |
|---|---|---|
| Manchu | 58 thousand letters | 1.2% |
| Latin (Cato) | 85 thousand letters | 3.1% |
| Occitan | 176 thousand letters | 12.6% |
| Mandarin | 314 thousand letters | 18.6% |
| Classical Japanese | 507 thousand letters | 27.3% |
| Nahuatl | 833 thousand letters | 32.0% |
| Middle Chinese | 1.59 million letters | 45.2% |
Eight corpora out of eight lined up strictly by file size. Manchu "survived" not because it is Manchu but because its text was the shortest we could find. Middle Chinese "failed" with the fattest.
A direct check confirmed it: one and the same Nahuatl, cut into pieces of different lengths, gives 2.1% twins at 60 thousand letters and 30.5% at 833 thousand. The language does not change - the thickness of the book does.
Why this happens. The larger the source text, the more varied the set of units from which the generator assembles its output, and the more rare words the output contains, words that occur two or three times. About a word met twice it is impossible to judge honestly what surroundings it lives in: a pair of chance neighbours will coincide with somebody else's pair purely by accident, and the program will count a pair of twins. The measure was gauging the poverty of the data, not the homophony of the cipher.
An honest comparison at equal size
We cut all eight corpora down to the same 57 thousand letters and ran them again.
| Language, equal size | Twins | Ratio R | Distance to the Voynich |
|---|---|---|---|
| Middle Chinese (Baxter) | 1.3% | 1.89 | 0.37 |
| Manchu (Banasik) | 1.4% | 2.02 | 0.47 |
| Latin (Cato) | 1.4% | 1.94 | 0.58 |
| Classical Japanese (bungo) | 1.9% | 1.78 | 0.84 |
| Occitan (Flamenca) | 1.7% | 1.68 | 1.01 |
| Mandarin (pinyin) | 1.7% | 1.88 | 1.04 |
| Sixteenth-century Nahuatl | 1.8% | 1.88 | 1.13 |
The spread collapsed from 0.43-37.2 to 0.37-1.24, every language landed inside the Voynich target or right up against it, and the order was inverted: Middle Chinese, which had been "the worst of all", turned out to be the closest.
Conclusion: the filter does not distinguish languages at all. It rejects no East Asian branch - and no other branch either.
What this means and what it does not
It does not mean that the Voynich manuscript is written in Chinese or in Manchu. It means exactly one thing: this method does not settle the question, and our earlier conclusion had no ground under it.
What is more, the result agrees with another measurement of ours, a far more important one. We tried to recover the cipher key from the inside, from the statistics of the text itself - and the attack failed: at the level of parts of words, the letter sequence of the original does not survive. The blindness of the language filter follows from that: if the trace of the letters has been erased, the language under the cipher does not show through either. More on this: why the Voynich manuscript cannot be read.
What really does argue against the Eastern version
We now use linguistic arguments neither for nor against. What remains is the non-linguistic ones, and they are untouched by our mistake:
- The parchment and the date. Radiocarbon gives 1404-1438, and the material is European parchment.
- The script. The letter shapes are European cursive of their period; a separate check showed that
this is not a Perso-Iranian or any other Eastern writing tradition. The forms of the signs are borrowed from the ordinary palaeography of the fifteenth century.
- The Occitan months. The zodiac pages carry month names written out, and they are Occitan, that is,
southern France. This is our anchor for localisation. The formula "the months are Piedmontese", which circulates in reviews, is wrong.
- The structure of the zodiac. The rings are built as a list of degrees, 30 to a sign - this is the
common Mediterranean astrological frame, not a Chinese speciality: no names of seasonal divisions and no Eastern constellations have been found in the text.
- The iconography. Chinese specialists who were asked about it found no Eastern motifs in the
drawings either.
Why we publish our own mistake
Because otherwise research turns into a collection of convenient results. We published the conclusion on 21-22 July and retracted it on 27 July - five days later, as soon as we saw the hidden variable. All the numbers and scripts are openly available, and the old results file has been left where it was with a retraction header on top: it is the trace of a mistake, not rubbish.
We wrote the general lesson into our working rules, and it is universal: a comparison must be set against a control that preserves the artefact of the procedure. Before saying "language X does not pass", run the same language at two or three different sizes and make sure the measure responds to the language and not to the thickness of the file. We had already had a similar lesson when testing the version about scribal abbreviations: there the gain came from the insertion procedure itself, not from the hypothesis.
No decipherment of the Voynich manuscript exists. We offer no readings - neither Eastern ones nor any others.
FAQ
Why is the Voynich manuscript linked to Asian languages at all?
Because of word length. The words in the manuscript are short and uniform in build, which is reminiscent of languages with a syllabic structure such as Chinese or Vietnamese. Hence Jacques Guy's suggestion of a monosyllabic language and Zbigniew Banasik's version of proto-Manchu.
So has the Chinese version been confirmed?
No. It is neither confirmed nor rejected - we simply have no working instrument with which to test it. We retracted our earlier conclusion that "the Eastern branches are filtered out" ourselves, once we found a hidden variable in the method.
What exactly was wrong with your test?
The key measure, the share of twin words, depended on the size of the source corpus rather than on the language. The corpora differed by a factor of 27, and the order of the results matched the order of the file sizes - eight out of eight. At equal size the difference between languages disappears.
What then argues against an Eastern origin of the manuscript?
Arguments that have nothing to do with language: the radiocarbon date of the European parchment, 1404-1438; European cursive as the basis of the letter shapes; Occitan month names in the margins of the zodiac; the structure of the zodiac rings as 30 degrees per sign; and the absence of Eastern iconography.