Language claims
"This is archaic Romani." What is left of the hypothesis after testing it against the manuscript's data
Romani groups appeared in Western Europe exactly when the manuscript's parchment was made. A fresh hypothesis about the language of the manuscript is built on that coincidence. We tested it against the transliteration, against the source record and against the statistics - and worked out what survived.

The idea looks attractive, and there is a real grain in it. Romani groups, who called themselves people out of "Little Egypt", appear in Western Europe in exactly the years to which radiocarbon assigns the parchment of the Voynich manuscript. A closed community, a language of its own that those around them could not understand, the trades of fortune telling and healing, a migration across the whole of Europe - and a book with an unread script, women's baths and herbs. The coincidence is too tempting to walk past.
In August 2026 a preprint appeared on Zenodo that turns this coincidence into a claim: the language of the manuscript is an archaic Balkan dialect of Romani, and the text is a herbal with water treatments and a section on women's health. We checked this work the way we check any other: against our own data, without indulgence and without prejudice.
What the hypothesis actually claims
The author takes about a dozen frequent words of the manuscript and offers a Romani etymology for each: daiin is read as daj "mother" and interpreted as "water", chey as čhaj "daughter" and interpreted as "fresh herb", qokal as kokalo "bone" and interpreted as "pounded root". Such words are said to cover up to eighty per cent of the tokens on individual pages, the Romani article and a case ending are said to be visible in the text, and thirty-seven pages of the manuscript are said to have been read already on this basis.
Check one: pages that do not exist
The first thing we did was to compare the list of "pages checked" against the manuscript itself. Five of the thirty-seven are absent from it: the leaves f60r, f61v, f62r, f74r and f110r are lost, and the transliteration has not a single line for them. A separate appendix to the work contains a "line-by-line translation of page f74r" - that is, a translation of a leaf that has not existed for several centuries.
This is not a quibble about presentation. Checking a reading begins with the existence of the text being read.
Check two: the lines do not match
Next we compared the "originals" quoted in the work against the transliteration. For page f85r the author quotes an opening line of five words. In the standard ZL transliteration, the one the whole community uses, the first line of the same locus consists of entirely different words - there is not a single match. What is being read, in other words, is not the manuscript but a retelling of it.
Check three: how much these words really cover
We counted the claimed coverage directly. Eleven of the proposed lexemes that do occur in the transliteration give 2462 occurrences out of the 37712 tokens of the manuscript - that is 6.53 per cent of the corpus. The maximum for a single page, among pages longer than fifty words, is 22.4 per cent. Eighty per cent is not reached anywhere.
Two further "lexemes" in the work's vocabulary do not belong to the EVA alphabet at all: they are entries in a different transliteration system, v101. One list mixes two incompatible ways of writing one and the same script - and the resulting "words" cannot be found in the text in either system.
One detail gives itself away: the "fifteenth-century recipes" come with millilitres and grams. The metric system was introduced in France in 1795.
What the source record says
Now about why the Romani version is difficult not only in this execution.
Romani remained unwritten until the twentieth century. The oldest known record of Romani speech is thirteen lines, about sixty words, in the manuscript collection of the Benedictine Johannes of Grafing, made presumably in Vienna around 1515. After that come the word lists of Andrew Borde of 1542, of Johan van Ewsum before 1570, of Bonaventura Vulcanius of 1597. Not a single text in Romani from the fifteenth century exists - no manuscript, no glossary, no marginal note.
Documents about the Roma in the fifteenth century, on the other hand, are many: Transylvania 1416-1418, the German lands 1417-1424, Switzerland, the Low Countries 1420, Bologna and Forlì in the summer of 1422, France 1419 and 1427, Aragon 1425. The groups carried letters of safe conduct and presented them - but they were not the ones who wrote them. The wording of the Aragonese archive is direct: this is a people that does not use writing to preserve its memory.
From this a simple consequence follows. To test the claim "the text is written in language X" one needs a corpus of language X from the same period. For Latin, for Italian, for Occitan such a corpus exists. For Romani of the fifteenth century it does not exist and cannot exist - the nearest record stands eighty years away from the upper bound of the radiocarbon date, and it is a word list of sixty words, not a text.
What the statistics say
We did nonetheless measure what can be measured. We took translations into three Romani dialects, truncated all the corpora to a common size of 15667 tokens - equal size is obligatory, otherwise the metrics measure the size of the file and not the language - and computed the conditional entropy per character.
| corpus | h2 |
|---|---|
| Voynich manuscript | 2.22 |
| Romani, Balkan dialect | 3.14 |
| Romani, Vlax dialect | 3.05 |
| Romani, Central dialect | 3.19 |
| Latin | 3.30 |
Romani falls exactly into the band where all natural languages lie, and stands almost a bit away from the manuscript. This is not an argument against the Romani line - it is the same argument that rules out Latin, Italian and all the other languages as the plaintext: between the language and the manuscript there necessarily stands a layer of encipherment. And once that layer is granted, languages stop being distinguishable: our own checks showed that after such a transformation the letter statistics of the source language do not survive at all.
What is left
Exactly one thing is left, and it is not the language. The coincidence in time and place is real: unlike the Aztec version, where the source is a century younger than the manuscript, here the overlap with the radiocarbon window is documented. So the meaningful question is not "what language is the book written in" but "are there traces of that milieu in the drawings": tents, wagons, characteristic headdresses, scenes of palm reading. That is tested by blind coding of features, the way we tested the chimeric quality of the plants - and it is tested honestly only with a criterion written down in advance.
The prior for such a check is low: outside researchers describe the costume and the architecture of the manuscript as Central European of the first half of the fifteenth century. But this is a question that can be answered with data rather than with a good story.
What this case teaches
The same thing the Lleida calendar and the "key" from the Max Planck Institute taught. The plausibility of a story and the testability of a claim are different things. A coincidence of dates is an invitation to a check, not its result. And checking a decipherment begins not with etymologies but with three boring questions: do the pages you are reading exist; do your lines match the transliteration; and how much of the text does your vocabulary really cover?
There is still no decipherment of the Voynich manuscript. A coincidence is not a reading.
FAQ
Could the Voynich manuscript have been written in Romani?
As a historical story - yes, the time and the place match: Romani groups are documented in northern Italy and the German lands in 1417-1422, and the parchment of the manuscript is radiocarbon-dated to 1404-1438. As a testable claim about language - no: Romani remained unwritten until the twentieth century, the oldest record of Romani speech dates from 1515, and there is simply nothing to compare the text of the manuscript with.
What is wrong with the published Romani decipherment?
Five of the thirty-seven "pages checked" in it do not physically exist - they are lost leaves of the manuscript, and one of the appendices translates a page that is not there. The lines of the "original" it quotes do not match the transliteration in a single word. The claimed coverage of "up to 80 per cent of the tokens" is in fact 6.53 per cent of the corpus.
The coincidence in time - surely that is a strong argument?
It is the best thing the version has, and it is true: by dates and geography the overlap is real, unlike the Aztec hypothesis, where the source is a century younger than the manuscript. But the presence of a people in the right place at the right time does not make a claim about language testable - a check needs a text, and no texts in Romani from the fifteenth century exist.
Does the statistics say anything about Romani?
We measured the conditional entropy of three dialects at equal corpus size: 3.05-3.19 bits against 2.22 for the Voynich. Romani falls into the ordinary band of natural languages and differs from the manuscript in exactly the same way as Latin. This is not an argument against the Romani line - it is a reminder that any version about language has to include a layer of encipherment, and once that layer is there languages become indistinguishable.