Language claims
"The key is found: it is vernacular Italian." What a numerical test of the Caspari and Faccini key showed
An archaeologist affiliated with the Max Planck Institute and an independent philologist read the star labels as POLAR and TOARE and declared the manuscript's language Italian. We reconstructed their key from their own examples and ran it over the whole manuscript against honest controls.

In February 2025 a preprint appeared with a confident title: "A Key to the Voynich Manuscript". The authors are Gino Caspari, an archaeologist affiliated with the Max Planck Institute and the University of Bern, and Agnese Faccini, an independent researcher. Their conclusion: the manuscript is written in vernacular Italian - the lingua volgare - recorded in medieval shorthand on top of a simple substitution cipher, and "the key to decipherment is now available".
The claim is more serious than the usual "decipherments" if only in its packaging: the preprint sits in the Max Planck Society's repository, and it comes with a key table and a vocabulary of more than 400 words. So we did what we do with every testable hypothesis: we reconstructed the authors' method exactly and ran it through controls that preserve all the freedoms of the method.
What the authors claim
The starting point is the labels next to stars in the astronomical section. Beside one of the stars the authors read the word POLAR - the Pole Star. From that anchor they unrolled an alphabet: some Voynich signs were declared Latin letters, some ligatures, and with this alphabet they read other labels: TOARE (Taurus, next to the Pleiades), O'CETE (the constellation Cetus), EXPE (Venus). They checked the star names against the printed Almagest of 1528.
Then comes the central assumption. The text, according to the authors, is written in shorthand: the scribe dropped vowels and cut off word endings. So the reading of a word is its "consonant skeleton", into which the reader is free to insert vowels and complete the ending. That is how a sequence of signs becomes, say, p(er)d(e)n or t(a)n(to): the letters in brackets are not in the manuscript.
The authors honestly call the key preliminary. But they state that with it "most words of the manuscript can be read", and that the language is confidently identified as vernacular Italian.
Their key can be tested rigorously - and we reconstructed it
The authors deliberately do not use machine-readable transliterations (EVA), so their table is a set of pictures of signs. For a numerical test we reconstructed the key in EVA notation not from the pictures but from their own addressed examples: the paper gives readings of specific lines on specific folios, and those lines are unambiguously locatable in the transliteration.
| Their reading | Word in the manuscript (EVA) | What it gives the key |
|---|---|---|
| torpore | dorchory | d=T, o=O, r=R, ch=P, y=E |
| p(er)d(e)n | chkar | k=D, ar=N |
| i(n) | s | s=I |
| fior | shor | sh=FI |
| o'c(e)fal | otchal | t=C |
| d(e)folie | ckholsy | ligature ckh=DF |
| exf(a)mo | ypchaiin | p=X, aiin=(a)MO |
The key reconstructed completely and is internally consistent: the same letter comes out identical from different examples. Then come three tests. The rules were fixed in advance, before the run.
Test one: the key against two hundred random keys
We take the whole manuscript - about 33 thousand words of running text. We decode every word with the authors' key and ask: does such a word exist in the Italian of their period? We built the lexicon from Dante and Boccaccio - the authors themselves name the Tuscan classics as the best matches for their vocabulary. The freedoms of the method are preserved in full: vowel insertion is allowed, so is truncation of the ending, so are the detachable article o' and the particle qo - everything the authors themselves use.
Then 200 random keys do exactly the same: the same signs, the same freedoms, but the letters shuffled at random.
The result: the authors' key really is better than a random one - on strict matching it reads 17.5% of words against 1.2% for a typical random key. Sounds like a win? No. The authors' key was tuned by hand so that frequent signs become frequent letters and the output looks Italian. Any key optimised for a language will beat a random one - that is a property of optimisation, not proof of a reading. The real questions are the other two.
Test two: does nonsense read as well?
If "reading most words" is the key's achievement, it should vanish when the language vanishes. We took the same Italian lexicon and shuffled the letters inside each word: the lengths and letter frequencies survived, the language itself was destroyed.
At the level of the authors' own rules (vowel insertion plus truncation) the key "reads" 69.7% of the manuscript's words against the real Italian lexicon. Against the shuffled pseudo-lexicon - 67.9%. The difference is less than two points. "Reading most words" is guaranteed by the freedom of the shorthand and does not depend on whether we are looking at Italian or at anagram noise.
Test three: why Italian, exactly?
The authors' main conclusion is the identification of the language. That can be tested directly: we give the same key with the same freedoms a lexicon of another language of the same period - Latin (Cato's agricultural treatise and the medical Antidotarium Nicolai), with the vocabulary sizes equalised.
Under the authors' own rules Latin reads BETTER than Italian: 90.3% against 65.8%. The method that "identified the language of the manuscript" cannot tell Italian from Latin - and in fact prefers Latin.
What is left of the hypothesis
Let us add up what the test showed, together with what is known about the manuscript independently of it.
- Even with vowel insertion, 62% of the manuscript's words do not read - "most words" is reached only at the rule level that reads nonsense just as well.
- Among the key's most frequent "readings" are forms like pue and dvte: not Italian words but old-orthography debris in the lexicon. The meaningful readings are rare islands, just as in every substitution "decipherment" before this one.
- The conditional entropy of the Voynich text is lower than that of every language tested, and simple substitution does not change entropy: an Italian text under such a cipher would remain statistically Italian. Our measurements point to a different class of mechanism - a verbose homophonic cipher.
- Voynich words are built as a rigid positional slot grammar - a real language with free shorthand never has that structure.
- The referent for the star names - the Almagest of 1528 - is a century later than the radiocarbon dating of the parchment, 1404-1438.
The conclusion: the reading is not confirmed. The Caspari and Faccini key is a frequency-aligned table optimised for an Italian-looking output. Its apparent success reproduces on a pseudo-language and is outdone by Latin under the authors' own rules.
What this case teaches
The requirement for any future "decipherment" follows from this test in a single line: show that your reading rules do NOT read nonsense. Until the freedom of a method (letter insertion, truncations, anagrams, choice among variants) has been measured on a control that preserves that freedom but kills the language, the count of "words read" means nothing. This is already the third hypothesis in our project to fall on exactly this control - and the first case where the test required reconstructing someone else's key from anchor examples. All the scripts and data of the check are open in the project's replication package.
FAQ
Did Caspari and Faccini decipher the Voynich manuscript?
No. Their own wording is cautious - a "preliminary key that may require substantial corrections". Our test showed that the claimed "reading of most words" reproduces just as well on a meaningless pseudo-lexicon, and that under their own rules Latin reads even better than Italian. The reading is not confirmed.
But their key produces meaningful Italian words - surely that is not a coincidence?
Not a coincidence - a property of the method. Their shorthand allows vowels to be inserted anywhere and word endings to be dropped. With that much freedom, almost any substitution table yields "readable" words. We measured this directly - a lexicon with the letters shuffled inside each word is read by their key almost as often as real Italian.
Why is the reference to the 1528 Almagest a problem?
The manuscript's parchment is radiocarbon-dated to 1404-1438. The authors check the star names against the Venetian edition of the Almagest printed in 1528 - a hundred years after the manuscript. To verify fifteenth-century names you need a fifteenth-century source; otherwise the match dates nothing.
Could the Voynich manuscript be a simple substitution cipher at all?
A measurable fact stands against it - the conditional entropy of the Voynich text is markedly lower than that of every European language tested. Simple letter substitution does not change a language's entropy, so it cannot turn an Italian text into a text with Voynich statistics. A more complex mechanism is needed - the data point to a verbose homophonic cipher.