Research results
What we know about the Voynich manuscript: the results in plain language
A summary of our entire project without formulas or jargon. What survived testing, what we rejected by the numbers, why decipherment is impossible without an external source - and why that is a result, not a defeat.

This page is a plain-language summary of the whole project: what we established, what we rejected, what remains open, and where we are heading. Every claim here rests on a numerical test; the detailed analyses are in the individual articles of this section, and the data and code are in the replication package on the Data page.
How we worked
Three rules hold the entire project together.
First: no claimed "readings". A coincidence is not a reading; any match with any language must pass a statistical test against chance.
Second: tests are registered in advance. The success threshold, the method, and the decision rule are fixed before the data are collected. There are no reruns "until the desired result"; failures are published alongside successes.
Third: blind coding. Wherever a feature is extracted from images, they were annotated by a person who did not know the hypothesis under test, working on shuffled, anonymized frames.
What is firmly established
The text is neither gibberish nor an imitation
The text shows Zipf's law, an alternation of "vowels" and "consonants", and a stable word grammar. The scenario of a scribe mindlessly copying himself was rejected by direct comparison: an imitation with matched statistics differs from the manuscript on five features out of six. Importantly, low entropy by itself does not rule out a hoax - that is the most common mistake in popular retellings.
The mechanism that produced the text has been reconstructed
This is the project's main result. The text was generated by a workshop technology: a syllabic cipher in which one syllable of the source language is replaced by a whole "word" drawn from substitution tables - about six such tables, shared by all five scribes of the manuscript. This single model accounts for all the famous oddities at once: the abnormally low predictability-entropy (2.1-2.2 bits against 2.96-3.23 for natural languages), the short words, the pairs of twin words, the difference between "languages A and B" (two lexicons of one system, not two languages), and the clustering of words by section. The model reproduces about ten independent statistics of the manuscript simultaneously - no other proposal comes anywhere close.
The captions next to the drawings are names, abbreviated into sigla
The short captions next to figures and plants behave like an inventory of names, not like connected speech: the share of unique words in them is 0.77 against 0.20 in the paragraphs. They are not numbers - we rejected that in four different ways. Each zodiac ring is its own list of roughly 20-25 names for 30 positions, and a caption itself is a two-letter abbreviation with a service marker meaning "this is a name". A siglum cannot be read without an external list of names - like initials without a phone book.
Origin: Europe, first half of the fifteenth century
Radiocarbon dating of the parchment gives 1404-1438, which rules out a late forgery. The month names written into the zodiac by a later hand are Occitan (southern France). The shapes of individual signs are ordinary European cursive of the period, but their function as abbreviations was tested and rejected: the scribes borrowed the look of the signs, not their meanings.
What is established locally or with reservations
The zodiac signal. An attribute of the "nymph" figures in the rings matched the table of masculine and feminine degrees by the astrologer al-Biruni at a shift of 23: a blind round gave 52 out of 78, p = 0.0088, with the peak exactly at the predicted shift. But we ourselves demonstrated its limit: the signal holds on three signs out of twelve and does not extend to the rest. Details are in the analysis of the zodiac degrees.
Counter-rotation of two rings. A blind recoding of 146 figures showed that on two inner rings of a single foldout the figures face against the general flow (10 of 11 and 4 of 5) - and these are the same signs that produced the zodiac signal. What this means is not yet known.
The "chimeric" plants. Plant features - leaf, root, flower - combine almost independently of one another in the manuscript, unlike in genuine herbals of the period: an association of 0.206 against 0.478 and 0.716 in two control herbals. The drawings are assembled combinatorially - just like the text.
The "inventory list" practice. The manuscript contains leaves where the scribe copies out a list of signs several times in a row, varying one or two positions. We compared three such leaves: they are different lists belonging to one scribal practice, not a single alphabet.
What was rejected by the numbers
The full catalogue is in the hypotheses section; here is the outcome in one table.
| Proposal | Test outcome |
|---|---|
| Readings via specific languages: Sanskrit, Slavic "bukvitsa", Hebrew anagrams, proto-Romance, Aztec, East Asian | Rejected structurally; beyond that, the popular "language filter" was shown not to distinguish languages at all |
| The Cardan grille and other mechanical generators | Rejected twice on different material |
| Mindless autocopying (a hoax) | Differs from the manuscript on five features out of six |
| Latin abbreviations, notarial and apothecary signs as a system | The shapes of the signs match, the functions do not; a blind test returned zero |
| Numbers anywhere: in the labels, in the text, in the figure counts | Three independent channels - three nulls |
| Church calendars as a key to the rings, including the Lleida calendar | Nine classes of sources - all nulls; the Lleida case passed three filters and turned out to be an artifact |
| Plant identification as a path to reading (the Bax path) | The plant name survives in none of the tested channels |
Why the manuscript cannot be read - and why that is a result
Decipherment requires recovering the key. We closed both paths to it - honestly and by the numbers.
The path from inside. We measured whether the letter sequence of the source text survives the encoding itself. The answer is no - the real text of the manuscript recovers worse than random noise. This is a property of the cipher itself: when one syllable can be written in many ways, the original letters are erased mathematically, and no computer can undo that.
The path from outside. One could try to find an external list - a calendar, a star catalogue, a roster of names - that matched the rings of the manuscript. We went through nine classes of such sources with a test that does not depend on the key. All nulls.
Hence the honest wording of the outcome: the mechanism of the text has been shown, the content cannot be read - and we showed why it cannot be read. This is a boundary established by measurement, not a complaint that "we failed". A verbose cipher without a bilingual document cannot be broken - just like a one-time pad, only for a different reason.
Where we are heading
Publication. Two scientific papers are finished and submitted: the first on the text mechanism, the second on the sigla-like labels and the negative testing programme. The replication package - all scripts, data, and reports - is already published on this site: anyone can reproduce every number of ours.
External leverage. The only thing that can move the case forward is new material: multispectral imaging of the remaining leaves of the manuscript (Yale has published ten so far), a new degree-by-degree list from the same tradition, or a bilingual document. Our tests are ready; checking any new candidate takes hours, not months.
What we will not do. Cycle through languages without a structural test, force numerology, or announce "readings". A coincidence is not a reading; this rule is what the project owes everything reliable in it. No decipherment is claimed.
FAQ
Has the Voynich manuscript been deciphered?
No. No confirmed decipherment exists - not by us and not by anyone else. Moreover, we showed why the text cannot be read without an external source - the encoding method does not preserve the letter sequence of the underlying language.
Is it a meaningful text or a hoax?
According to our tests, the text was produced by a purposeful technology - a syllabic cipher with substitution tables shared by several scribes. The mindless-copying scenario was rejected by direct comparison, but not on the grounds of low entropy - low entropy by itself settles nothing.
What might the manuscript be about?
The working picture is a calendrical-astrological handbook with a women's-medicine focus, made in the first half of the fifteenth century, probably in southern France. This is a hypothesis grounded in structure, not a reading of the text.
What could change the situation?
An external source - a bilingual document, a list of names from the same tradition, or new data such as multispectral imaging of the remaining leaves. Our scripts are ready to test any new candidate within hours.
Where can I see the data and code?
The replication package with scripts, data, and reports is published on this site in the Data section. Both scientific papers have been submitted for publication.