From cuneiform to Linear Elamite to the Dhofari script, learn how archaeologists decipher ancient writing systems and which scripts remain unsolved in 2026.
Every few years, archaeologists announce they have deciphered an ancient script. Linear B took 50 years. Maya glyphs took 40. The Indus Valley script is still unsolved. In 2025, the Dhofari script of Oman was deciphered for the first time. In 2026, a French researcher cracked Linear Elamite, a 4,000-year-old script from Iran.
The methods behind these breakthroughs are remarkably consistent. They rely on bilingual texts, proper names, frequency analysis, and increasingly, machine learning. The story of each decipherment is also a story of the people who spent decades on work that might never pay off.
Every ancient script that has ever been deciphered needed an anchor. Usually that anchor is a bilingual text, like the Rosetta Stone, which presents the same content in both a known and an unknown script. Sometimes the anchor is a known language family: if the unknown script encodes a language related to one we already understand, phonetic values can be guessed and tested.
Bilingual texts. The Rosetta Stone (196 BCE) presented the same decree in hieroglyphics, Demotic, and Greek. Champollion used the Greek to identify proper names (Ptolemy, Cleopatra) in the hieroglyphic text, then worked outward from those phonetic values. The same principle was used for cuneiform: the Behistun Inscription, a trilingual text in Old Persian, Elamite, and Babylonian, was the key that unlocked Mesopotamian writing.
Proper names. Names of kings, gods, and places are the easiest entry point because they are phonetic, not ideographic. Even in a script that uses logograms (signs representing whole words), names are usually spelled out syllable by syllable. As Francois Desset explained about his decipherment of Linear Elamite in 2026: "The key to deciphering a script, as is so often the case, lies in proper names: names of places, gods, kings."
Frequency analysis. The same statistical methods that break classical ciphers apply to ancient scripts. Counting sign frequencies, identifying repeated sequences, and analyzing word lengths all provide clues. Alice Kober used exactly these methods in her work on Linear B in the 1940s, cataloging 180,000 index cards of sign data by hand.
Machine learning. As of 2026, AI is being applied to undeciphered scripts, but with mixed results. AI excels at large-scale pattern testing: it can check whether a hypothesis about one sign holds up across an entire corpus in seconds. But it cannot manufacture meaning from nothing. It needs an anchor (a known language family or bilingual text) to tell it which patterns are meaningful. A June 2026 claim that Linear A represents a Semitic language, made by an amateur researcher using AI scripts, remains under review and is not accepted by the scholarly community.
Cuneiform is the oldest known writing system, developed by the Sumerians in Mesopotamia around 3200 BCE. It was written by pressing a reed stylus into clay tablets, producing wedge-shaped marks. The script was used for over 3,000 years across multiple languages (Sumerian, Akkadian, Babylonian, Assyrian, Old Persian).
The decipherment began with Georg Friedrich Grotefend in 1802. Grotefend, a German schoolteacher, focused on the Behistun Inscription, a massive trilingual text carved into a cliff in Iran by Darius the Great. He guessed that certain repeated sequences in the Old Persian version were royal names (Darius, Xerxes) and identified the phonetic values of several signs by matching the names against known historical records.
Henry Rawlinson, a British East India Company army officer, extended Grotefend's work in the 1830s and 1840s. He risked his life climbing the Behistun cliff to copy the inscription, then used the trilingual structure to decode the Babylonian and Elamite versions. By the 1850s, cuneiform was substantially deciphered, opening up three millennia of Mesopotamian history.
Linear B, the script of Mycenaean Greece (1450-1200 BCE), was deciphered in 1952 by Michael Ventris, an architect who had worked as a codebreaker at Bletchley Park during World War II. The decipherment was built on the work of Alice Kober, who identified grammatical inflection in the script and constructed a phonetic grid before her death in 1950.
Ventris's key insight was assuming Linear B encoded Greek, a hypothesis that almost no other scholar held. He tested the assumption against the Pylos and Knossos tablets and found that the resulting words made sense in an archaic Greek dialect. The decipherment, published with John Chadwick in 1953, pushed the documented history of the Greek language back by centuries.
A 2024 donation of papers by Thomas Palaima to the University of Cincinnati archives has provided new material for studying the decipherment process. A 2024 updated edition of Documents in Mycenaean Greek, covering 50 years of subsequent scholarship, was also published. New tablets continue to be discovered at sites including Thebes, Khania, and Ayios Vasileios.
Maya hieroglyphic writing, used in Mesoamerica from approximately 300 BCE to the 16th century CE, was deciphered over several decades, with the breakthrough coming from Yuri Knorosov, a Soviet epigrapher, in the 1950s. Knorosov worked from the Dresden Codex, one of only four surviving Maya bark-paper books, and a Spanish bishop's 16th-century manuscript that included a "Maya alphabet" (actually syllabic signs).
Knorosov realized that Maya glyphs represented syllables, not just ideograms. He identified phonetic values for several signs and showed that the script encoded a Mayan language (Yucatec). His work was published in 1952 but was largely ignored in the West due to Cold War tensions. It was not until the 1970s and 1980s that his approach was widely adopted, and the 1990s before most Maya texts could be read fluently.
The decipherment revealed that Maya inscriptions were primarily historical records: dynastic histories, military victories, astronomical observations, and religious ceremonies. This overturned the earlier assumption that the Maya were a peaceful civilization of astronomers.
Dhofari script (2025). In July 2025, Ahmad Al-Jallad, a linguist at Ohio State University, deciphered the main subtype of the Dhofari script, found on rock faces in Oman and Yemen. The script is approximately 2,400 years old and had defied decipherment for over a century. Al-Jallad identified an abecedary (a listing of the alphabet's letters in order) in a cave inscription, then matched the Dhofari letter sequence to the halham order used in ancient South Arabian scripts. This allowed him to assign sound values to the letters and begin reading words of a previously unknown ancient Arabian language.
Linear Elamite (2026). In April 2026, French archaeologist Francois Desset deciphered Linear Elamite, a 4,000-year-old script from Bronze Age Elam (modern-day Iran). The script, made up of 77 geometric signs, had been rediscovered in 1903 but stumped experts for over a century. Desset's breakthrough came from access to ten new texts on vases from the Mahboubian collection in London. He identified the name of a ruler, Shilhaha (reigned around 1950 BCE), by noticing a repeated ending pattern in a four-symbol sequence. From that single proper name, he expanded to 45 readable inscriptions. His work has been compared to Champollion's decipherment of Egyptian hieroglyphics.
Sidetic alphabet (2026). In June 2026, archaeologists in Side, Turkey, identified five additional letters in the Sidetic alphabet, bringing the total to 31. Sidetic is a lost language related to Lycian and Carian, once spoken in the ancient port city of Side in Anatolia. The letters were found in bilingual inscriptions of 30-40 lines, a substantial find given that most Sidetic inscriptions are only one or two lines long.
Herculaneum scrolls (2026). In June 2026, the Vesuvius Challenge announced the complete virtual unwrapping and reading of PHerc. 1667, a carbonized papyrus scroll from Herculaneum, buried by the eruption of Mount Vesuvius in 79 CE. This is the first Herculaneum scroll to be fully digitally unrolled and read. The scroll contains a previously unknown text by the Epicurean philosopher Philodemus, titled On Gods, Book 8. The technology uses high-resolution X-ray microtomography and machine learning to detect ink inside carbonized papyrus without physically opening it. A third scroll, PHerc. 139, had its title and author attribution recovered, identifying it as another work by Philodemus.
Several ancient scripts remain unreadable as of 2026:
Linear A. The predecessor to Linear B, used by the Minoan civilization on Crete (1800-1450 BCE). The surviving corpus is only about 7,500 characters. The language it encodes has no confirmed relationship to any known language. It is a language isolate, which means there is no anchor for decipherment. AI-assisted attempts have been made but none are accepted by scholars.
Indus Valley script. Used by the Indus Valley Civilization (2600-1900 BCE) in what is now Pakistan and India. The surviving corpus consists of very short inscriptions (mostly 5-10 symbols) on seals and tablets. The brevity of the texts makes statistical analysis nearly impossible. The language is unknown, and there is no bilingual text. Some researchers question whether the symbols constitute a writing system at all.
Rongorongo. The script of Easter Island (Rapa Nui). It is possibly one of the few independent inventions of writing in human history. Most of the wooden tablets bearing the script were destroyed by European colonizers in the 19th century. The surviving corpus is small, and the language (Rapa Nui) is known but the script's structure is not understood well enough to map signs to sounds.
Etruscan. The language of pre-Roman Italy. A partial vocabulary has been collected from short funerary inscriptions, and the alphabet is readable (it is derived from Greek). But the grammar and deeper meaning of longer texts remain elusive because there is no bilingual text of sufficient length.
Decipherment requires an anchor. Without a bilingual text or a known language family, no amount of computational power can produce a reliable translation. AI can test hypotheses at scale, but it cannot invent the initial hypothesis from nothing. The Conversation's 2026 analysis of AI and decipherment concludes that AI is a powerful research assistant for cross-checking human hypotheses against large corpora, but statistical pattern matching alone cannot manufacture meaning.
Short corpora are the other hard limit. The Indus Valley script has approximately 3,700 known inscriptions, but most are under 10 symbols. Linear A has about 7,500 characters total. These sample sizes are too small for the statistical methods that cracked Linear B (which had thousands of tablets) or cuneiform (which had the massive Behistun Inscription).
As of 2026, the most recent decipherments are Linear Elamite by Francois Desset (announced April 2026), the Dhofari script by Ahmad Al-Jallad (July 2025), and the complete virtual reading of a Herculaneum scroll by the Vesuvius Challenge (June 2026). The Sidetic alphabet also had five new letters identified in June 2026.
The major undeciphered scripts are Linear A (Minoan Crete, ~1800-1450 BCE), the Indus Valley script (~2600-1900 BCE), Rongorongo (Easter Island), and Etruscan (pre-Roman Italy, partially readable but grammar is unclear). The Indus Valley script is particularly challenging because the inscriptions are very short and the language is unknown.
AI can help by testing hypotheses at scale, spotting patterns humans might miss, and restoring damaged inscriptions. But it cannot decipher a script on its own. It needs an anchor: a known language family or a bilingual text. AI-assisted claims about undeciphered scripts (like Linear A) remain under peer review and are not accepted by the scholarly community as of 2026.
Jean-Francois Champollion used the Rosetta Stone, a trilingual text in hieroglyphics, Demotic, and Greek. He identified royal names (Ptolemy, Cleopatra) in the Greek text, then located the same names in cartouches (oval frames) in the hieroglyphic text. From those phonetic values, he expanded to other signs. He published his decipherment in 1822 in the Lettre a M. Dacier.
In June 2026, the Vesuvius Challenge announced the complete virtual unwrapping and reading of PHerc. 1667, a carbonized papyrus scroll from Herculaneum buried by Mount Vesuvius in 79 CE. Using high-resolution X-ray microtomography and machine learning, researchers read the scroll without physically opening it. The text is a previously unknown work by the Epicurean philosopher Philodemus, titled On Gods, Book 8.
Code Identifier
Identify the cipher or encoding used in a piece of text. Paste encoded or encrypted data and the code identifier returns ranked candidates with confidence scores.
Hieroglyphics Translator
Convert English text to Egyptian hieroglyphs using Unicode uniliteral signs from the Gardiner Sign List, and reverse hieroglyphs back to text.
Linear B Script Translator
Convert Latin transliteration to Linear B syllabograms (Unicode U+10000) and back. The oldest known form of Greek, used 1450-1200 BCE.
Letter Frequency Analyzer
Count and analyze letter frequencies in text for cryptogram solving.
Linear B: How a 3,000-Year-Old Script Was Finally Deciphered
Michael Ventris deciphered Linear B in 1952, proving it was an early form of Greek. Learn how Alice Kober's work made it possible and how new tablets are still being discovered.
Frequency Analysis Explained: How to Break Any Substitution Cipher
Al-Kindi discovered frequency analysis in 9th-century Baghdad. The technique still breaks CTF substitution ciphers today. Here is how it works and how to apply it.