A 240-page book in a script that matches no known language, with plants that do not exist, carbon-dated to the early 1400s. Studied for 600 years. Still unread. Here is what we know and what we do not.
A book nobody can read, in a script that matches no known language, with illustrations of plants that do not exist. It has been studied for over 600 years. It remains a mystery.
The Voynich Manuscript is approximately 240 pages of vellum, written in an unknown script with approximately 38,000 words distributed across roughly 25,000 glyphs. The pages are filled with botanical, astronomical, balneological, and pharmaceutical illustrations. Carbon dating of the vellum places it between 1404 and 1438 CE. The ink and paint are consistent with medieval materials.
No one has produced a decipherment that the academic community accepts. Every few years, someone announces they have cracked it. None of these claims have survived scrutiny. The manuscript sits at Yale University's Beinecke Rare Book and Manuscript Library as MS 408, and the full digitized manuscript is available online.
You can apply the same analytical tools used on the Voynich text with our Letter Frequency Analyzer and Cipher Identifier.
The Voynich Manuscript was discovered by the Polish-American book dealer Wilfrid Voynich in 1912 (or possibly 1911) in a Jesuit college near Rome. It measures 23.5 by 16.2 cm and contains approximately 240 vellum pages. Based on the vellum's thickness and the number of missing pages, the original manuscript likely had about 272 pages.
The content is divided into sections based on the illustrations:
Botanical section: Hand-drawn plants, most of which do not match any known species. Each plant is accompanied by text in the unknown script.
Astronomical section: Zodiac symbols, suns, moons, and stars. Some pages contain circular diagrams with figures in pipes or tubes.
Balneological section: Figures (mostly nude women) in pools, baths, and interconnected tubes. This section has generated the most speculation, as the imagery is bizarre and does not match known medieval medical texts.
Pharmaceutical section: Drawings of roots, leaves, and what appear to be apothecary jars.
Recipes section: Short paragraphs, possibly recipes or incantations, each marked with a star or asterisk-like symbol.
Radiocarbon dating conducted at the University of Arizona in 2009 placed the vellum between 1404 and 1438 CE with 95% confidence. The ink, paint, and pigments have been analyzed and are consistent with medieval European materials. The manuscript is authentic to the early 15th century.
The Voynich text has been subjected to extensive statistical analysis. These analyses reveal properties that are both language-like and unlike any known language or cipher.
Zipf's Law: The word frequencies in the Voynich text follow Zipf's Law, a statistical distribution where the frequency of a word is inversely proportional to its rank. This property is found in natural languages and is generally absent in random gibberish. The presence of Zipf's Law suggests the text encodes real information, though it does not prove it.
Word length distribution: Voynich words tend to be short (3 to 10 glyphs) with a distribution that resembles natural languages. However, the glyph combinations are highly constrained. Certain glyphs appear almost exclusively at the beginnings of words, others at the ends, and some never appear in certain positions. This positional restriction is characteristic of real writing systems.
Repeated words: The text contains unusually frequent word repetitions. In some passages, the same word appears three or more times in a row, which is rare in natural languages but can occur in certain types of enciphered text or shorthand.
Low character entropy: The conditional entropy of Voynich text (how predictable the next glyph is given the previous glyphs) is lower than most natural languages. This could indicate a verbose cipher (where single plaintext letters are encoded as multiple glyphs) or a shorthand system.
In November 2025, science journalist Michael Greshko published a study in the journal Cryptologia describing the "Naibbe cipher," a homophonic substitution cipher that produces ciphertext statistically similar to the Voynich text. The study demonstrated that a 15th-century scribe could have produced Voynich-like text using a cipher that encrypts Latin or Italian. This does not prove the Voynich Manuscript is enciphered, but it shows the ciphertext hypothesis remains viable. A February 2025 preprint on OSF proposed that the manuscript is vernacular Italian written in shorthand, identifying over 400 words, though this claim has not been independently verified.
The manuscript has defeated every codebreaker who has attempted it.
William Friedman (1891-1969), the chief cryptanalyst of the US National Security Agency's predecessor, spent approximately 30 years on the Voynich Manuscript in his spare time. He tried every classical cipher technique available and failed. He eventually hypothesized it might be a constructed artificial language, but he never produced a translation.
Stephen Bax (2014) claimed to have deciphered 14 words and 10 glyphs by identifying proper nouns (plant names) in the botanical section. His method was similar to Champollion's approach to hieroglyphics: find a proper noun, match it to the text, derive phonetic values. His claims were disputed and have not been accepted by the academic community.
University of Alberta AI study (2016): Computer scientists used machine learning to test the hypothesis that the underlying language is Hebrew. They claimed 80% of their decoded words matched Hebrew dictionaries. The study was widely criticized for methodological flaws, including the fact that the algorithm was told to look for Hebrew and could produce similar results for other languages.
Raymond Aranyics (2026) published a preprint on OSF claiming to decode the manuscript using the Latin Vulgate Bible as a steganographic grid, with a chapter-and-initial-letter algorithm applied across 170,000 characters. Like previous claims, this has not been independently verified.
The pattern is consistent: a researcher announces a decipherment, the claim generates media attention, other experts find flaws, and the manuscript remains unread. The fundamental problem is that without a bilingual text (like the Rosetta Stone) or a known plaintext, any proposed decipherment is unfalsifiable. You can always find a mapping that produces plausible output from a sufficiently flexible system.
Three theories dominate the discussion:
Cipher hypothesis: The text is encrypted using a cipher that has not yet been identified. The statistical properties (Zipf's Law, positional glyph restrictions) support this. Greshko's 2025 Naibbe cipher study shows that a historically plausible cipher could produce Voynich-like text from Latin or Italian input. The challenge is that no one has produced a consistent decryption that produces readable text across the entire manuscript.
Constructed language hypothesis: The text is written in an artificial language invented by the manuscript's author. Friedman eventually leaned toward this theory. A constructed language could explain the unusual statistical properties (low entropy, strict positional rules) without requiring a cipher. The problem is that constructed languages from the medieval period are extremely rare, and none match the Voynich script.
Hoax hypothesis: The text is meaningless gibberish generated to look like writing, possibly to sell the manuscript as a valuable curiosity. Gordon Rugg argued in 2004 that the text could be produced using a Cardan grille, a steganographic tool that generates pseudo-text with language-like statistical properties. However, the Zipf's Law distribution and the consistency of glyph patterns across the entire manuscript make a hoax less likely, as generating consistent gibberish across 240 pages by hand would be remarkably tedious.
The academic consensus, as of 2026, is that the manuscript remains undeciphered. No proposed solution has achieved widespread acceptance.
The Voynich Manuscript is a case study in the limits of cryptanalysis without context. Our Cipher Identifier can identify common ciphers (Caesar, Vigenere, substitution, Base64, etc.) because it has known patterns to match against. The Voynich script matches no known cipher, language, or encoding system.
The lesson is that statistical analysis alone is insufficient for decipherment. You need either a known-plaintext equivalent (like the Rosetta Stone) or a known cipher system with a recoverable key. The Voynich has neither. Until someone finds a bilingual text or a documented cipher system that produces Voynich-like output, the manuscript will likely remain unread.
You can analyze the Voynich text yourself using our Letter Frequency Analyzer to see the Zipf's Law distribution and positional glyph restrictions that have puzzled codebreakers for a century.
No. Despite numerous claims over the past century, including studies in 2025 and 2026 proposing Italian shorthand and Latin Vulgate steganographic solutions, no decipherment has been accepted by the academic community. The manuscript remains undeciphered as of 2026.
Radiocarbon dating of the vellum places it between 1404 and 1438 CE with 95% confidence. The ink and pigments are consistent with medieval European materials. The manuscript is authentic to the early 15th century.
The author is unknown. The manuscript's history can be traced to the early 17th century when it was in the library of the Holy Roman Emperor Rudolf II in Prague, who reportedly purchased it for a large sum. Earlier provenance is uncertain.
The underlying language (if any) is unknown. Proposed languages include Latin, Italian, Hebrew, and various constructed languages. A 2025 study in Cryptologia demonstrated that a cipher could produce Voynich-like text from Latin or Italian, but this does not confirm the underlying language.
It is held at Yale University's Beinecke Rare Book and Manuscript Library as MS 408. The full manuscript has been digitized and is freely available online through the Beinecke Library's website.
Code Identifier
Identify the cipher or encoding used in a piece of text. Paste encoded or encrypted data and the code identifier returns ranked candidates with confidence scores.
Letter Frequency Analyzer
Count and analyze letter frequencies in text for cryptogram solving.
Cryptogram Solver
Automated solving of substitution ciphers using frequency analysis and pattern recognition.
The Rosetta Stone: How One Rock Deciphered Ancient Egyptian Hieroglyphics
For 1,400 years nobody could read hieroglyphics. Then a French soldier found a trilingual stone in 1799, and Jean-Francois Champollion spent 20 years cracking it. Here is how he did it.
Frequency Analysis Explained: How to Break Any Substitution Cipher
Al-Kindi discovered frequency analysis in 9th-century Baghdad. The technique still breaks CTF substitution ciphers today. Here is how it works and how to apply it.