Language claims
Cheshire's proto-Romance version: a decipherment that cannot be tested
In 2019 a university press release announced that the manuscript had been read in a fortnight. A few days later the release was taken down. We look not at personalities but at the construction: why a version of this kind is untestable in principle.

In 2019 this story travelled far beyond a narrow circle: a university press release announced that the Voynich manuscript had been read, and that the work had taken about a fortnight. A few days later the release was removed from the website, and specialists in Romance philology and in the manuscript itself reacted sharply against it.
What we take apart here is not personalities and not emotions but the construction of the version: why it is built in such a way that it can be neither confirmed nor refuted.
What the claim is
In brief: the text is written in "proto-Romance", a lost spoken language, ancestor of Italian, Spanish, Portuguese and the rest; the signs are read directly and there is no cipher at all; find the sound values and the book can be read.
It sounds economical. The trouble starts at the next step.
Freedom of substitution kills testability
In the system as presented, one and the same sign may stand for different sounds, and the choice is made separately for each word, on the basis of what seems to make sense in that place.
Consider what that means. Suppose each sign has at least three admissible readings. Then a word of five signs yields about two hundred and forty combinations, and out of that set there will nearly always be something resembling a word of a living Romance language - especially if the comparison is made not against one language but against a whole family plus a reconstruction.
Hence a consequence that matters more than any particular quibble: a system like this will read anything. Give it a random set of signs of the same length and it will produce a text no less meaningful. Which means that a successful reading proves nothing, and that there is nothing with which to refute it either.
A scientific hypothesis is obliged to forbid some outcomes. This one forbids none.
What is missing: a prediction
What would a testable version of the same class look like? Roughly this:
- fix the sign-to-sound table before translating, and never change it afterwards;
- present a reading of a stretch of text that took no part in the fitting;
- show that the grammar comes out regular: the same endings in the same positions;
- predict something independent - for instance, that a particular word will occur on the page with a
particular drawing - and check it.
Not one of these steps was taken in the work. That is precisely why specialists say not "a mistake in the details" but "this is not a testable claim".
What our measurements say
Here we can add something of our own, and it has nothing to do with taste.
The version requires that the text be read directly: sign for sound, word for word. That is the possibility we tested - we tried to recover the key from the internal statistics of the manuscript. The attack failed with a telling result: the real text came out worse than random noise in terms of recoverability. The letter sequence of the original does not survive at the level of parts of words.
That is incompatible with direct reading. If sounds of a language stood behind the signs, the trace of the language would show through in the statistics of combinations - and it does not.
The second observation concerns the structure of the word. Three quarters of the words in the manuscript are assembled on a rigid "prefix plus root plus suffix" scheme out of a very small set of elements: about 56 roots cover 80% of the text. For comparison, Latin needs about 640 roots for the same 80%. A natural language is not built that way - this is the fingerprint of a table of codes, not of the vocabulary of living speech.
Why versions like this appear regularly
This is the general trap of the field. The manuscript holds about 37 thousand words, the syllables of any language combine in a limited number of ways, and at that volume any language will yield dozens of handsome coincidences. That is how the Persian, Hebrew, Sanskrit and Aztec readings were born. Every author sincerely sees his own.
The only defence is the order of operations: structure first, meaning afterwards. First the measurable prediction, then the translation. If a version begins with the translation, there is no way to test it, however convincing the resulting text may look.
That is also why we did not "run Cheshire" through our filter: what can be run is a hypothesis that makes a prediction. Here there is no prediction.
No decipherment of the Voynich manuscript exists - neither ours nor anybody else's. A coincidence is not a reading.
FAQ
What did Gerard Cheshire claim?
That the Voynich manuscript is written in "proto-Romance", a lost spoken ancestor of the Romance languages, that the script contains no unusual signs, and that the text can be read directly, without a key. The paper appeared in 2019, and the university press release about it was withdrawn shortly afterwards.
Why did specialists reject it?
The main objection is methodological: in the proposed system the same sign may be read as different sounds, and the choice is made separately for each word. With that much freedom, a meaningful phrase can be obtained from almost any set of signs, and there is nothing left with which to test the claim.
Does a "proto-Romance" language exist?
As a reconstruction of the popular spoken language, yes - specialists work with Vulgar Latin. But there is no single written proto-Romance corpus against which a reading could be checked; in Cheshire's version it is effectively constructed as the translation proceeds.
What do your own measurements say?
That direct reading is impossible in principle. We tried to recover the key from the statistics of the text itself, and the real text came out worse than random noise. The letter sequence of the original does not survive at the level of parts of words, which means there is nothing there to "read directly".