Al-Kindi discovered frequency analysis in 9th-century Baghdad. The technique still breaks CTF substitution ciphers today. Here is how it works and how to apply it.
Al-Kindi solved encrypted Arabic manuscripts in Baghdad around 850 CE using a technique that still breaks CTF challenges today. He documented it in Risalah fi Istikhraj al-Mu'amma (A Manuscript on Deciphering Cryptographic Messages), the earliest known work on cryptanalysis and the first text to describe frequency analysis as a systematic attack.
The technique does not require algebra, computing, or advanced mathematics. It requires counting. Paste any monoalphabetic substitution ciphertext into the Letter Frequency Analyzer and the distribution will make the shift obvious within seconds.
This post covers the theory, a worked Caesar-3 example, and the practical limits of the method. If you want to follow along with real ciphertext, keep the analyzer open in a second tab.
A monoalphabetic substitution cipher replaces each letter with exactly one other letter, consistently throughout the message. The Caesar cipher is the simplest case: every A becomes D, every B becomes E, with a fixed shift of 3. A more general substitution cipher uses a random permutation of the alphabet. A might become Q, B might become T, Z might become F. The mapping is arbitrary, but each plaintext letter always maps to the same ciphertext letter.
That consistency is the weakness. Natural language is not random. In English text of any reasonable length, the letter E appears approximately 12.7% of the time. T appears around 9.1%, A around 8.2%, O around 7.5%. These proportions are stable across most written English, regardless of topic, author, or century.
When you apply a monoalphabetic substitution, you shift or permute the alphabet but you do not change the frequency distribution. You just relabel it. The most common letter in the ciphertext still corresponds to E in the plaintext. The second most common corresponds to T. The shape of the distribution survives the substitution intact.
The Wikipedia letter frequency table and Peter Norvig's corpus analysis at norvig.com/mayzner.html both confirm this distribution across large English corpora. Norvig's work, based on Google Books Ngram data covering trillions of words, matches the older manual counts from Mayzner's 1965 study to within a fraction of a percent. The distribution is not a curiosity. It is a statistical fingerprint of the language.
Letter frequencies in English
The standard frequency order for English is approximately:
``
E T A O I N S H R D L C U M W F G Y P B V K J X Q Z
``
The mnemonic ETAOIN SHRDLU covers the twelve most common letters. This order was so well known in the typesetting industry that Linotype machines arranged their keys in frequency order. ETAOIN SHRDLU is literally the first two rows of the Linotype keyboard, left to right. Operators who needed filler text would run their fingers down the first two columns, which is why the nonsense string "ETAOIN SHRDLU" occasionally appeared in print as a typesetter's error.
The approximate percentages for the top letters:
| Letter | Frequency | Letter | Frequency |
|---|---|---|---|
| E | 12.7% | M | 2.4% |
| T | 9.1% | W | 2.4% |
| A | 8.2% | F | 2.2% |
| O | 7.5% | G | 2.0% |
| I | 7.0% | Y | 2.0% |
| N | 6.7% | P | 1.9% |
| S | 6.3% | B | 1.5% |
| H | 6.1% | V | 1.0% |
| R | 6.0% | K | 0.8% |
| D | 4.3% | J | 0.2% |
| L | 4.0% | X | 0.2% |
| C | 2.8% | Q | 0.1% |
| U | 2.8% | Z | 0.1% |
Bigrams and trigrams
Single-letter frequency is a first-pass attack. Bigrams (two-letter pairs) and trigrams sharpen the analysis considerably. The most common English bigrams are TH, HE, IN, ER, AN, RE, ON, EN, AT. The most common trigrams are THE, AND, ING, HER, HAT, HIS, THA, ERE.
If a ciphertext trigram appears far more often than any other three-letter group, it is almost certainly THE. THE is roughly twice as frequent as the next trigram (AND), which makes it a reliable anchor. Once you have THE, you immediately know three letter mappings, and those three letters appear in dozens of other common words.
The 7-step attack procedure
1. Count the frequency of every letter in the ciphertext. 2. Rank them from most to least frequent. 3. Map the most frequent ciphertext letter to E, the second to T, and so on down the ETAOIN SHRDLU order. 4. Partially decrypt the ciphertext with these initial guesses. 5. Look for nearly-complete words. A three-letter sequence where two letters are known often reveals the third by word shape alone. 6. Adjust mappings where partial words do not make sense. If your guessed E produces a word like QEZ, something is wrong. 7. Repeat until full plaintext is recovered.
The Substitution Cipher Helper lets you work through this interactively, swapping letter assignments as you go. You can also use the Cryptogram Solver to automate the statistical mapping for standard puzzles.
Consider this ciphertext, produced by a Caesar-3 shift:
``
WKHU LV D WUDFH RI WKH RULJLQDO PHVVDJH KLGGHQ
LQ WKH SDWWHUQ RI OHWWHUV. IUHTXHQFB DQDOBVLV
ZLOO UHYHDO LW. WKH PRVW FRPPRQ OHWWHU KHUH LV H.
``
Step 1: Count frequencies. H appears most often (14 times). W, K, and V are also high. Z does not appear at all.
Step 2: Map H to E. H is the most frequent ciphertext letter, and E is the most frequent English letter at 12.7%. Assign H to E.
Step 3: Look for single-letter words. D appears alone twice. In English, single-letter words are either A or I. Try D to A first.
Step 4: Examine the trigram WKH. It appears three times. With H mapped to E, WKH is now `_ _ E`. The most common English three-letter word ending in E is THE. Assign W to T and K to H.
Step 5: Check WKH LV. With W and K resolved, WKH LV becomes THE `_ _`. The two-letter word IS fits the context. Assign L to I and V to S.
At this point the substitutions T, H, E, A, I, S are in place:
``
THE_ IS A T_A_E O_ THE O_I_I_AL _E__A_E HI__E_
I_ THE _ATTE_I O_ _ETTE_S. I_E_E__C_ A_AL_SIS
_I__ _E_EA_ IT. THE _OST __O__O_ _ETTE_ HE_E IS E.
``
The remaining letters fall out from word patterns. T_A_E is TRACE, so U to R and F to C. _E__A_E is MESSAGE, so P to M and V (already S) confirms. The full plaintext reads:
``
THERE IS A TRACE OF THE ORIGINAL MESSAGE HIDDEN
IN THE PATTERN OF LETTERS. FREQUENCY ANALYSIS
WILL REVEAL IT. THE MOST COMMON LETTER HERE IS E.
``
This was a Caesar-3 ciphertext, which you could also break by brute force (only 25 possible shifts). But frequency analysis works equally well on random permutation ciphers where brute force is not feasible. A random monoalphabetic substitution has 26 factorial possible keys, roughly 4 times 10 to the 26th power. Exhaustive search is hopeless. Frequency analysis still breaks it in minutes.
CTF cryptogram challenges: The "Crypto" category on platforms like picoCTF and CryptoHack routinely includes monoalphabetic substitution puzzles. The expected solver chain is: paste ciphertext, run frequency analysis, build a letter mapping, verify against word patterns. Most of these challenges are solvable in under five minutes once you know the procedure.
Newspaper cryptograms: Published cryptograms, including American Cryptogram Association puzzles and newspaper daily puzzles, are all monoalphabetic substitutions. The solving technique is the same as the CTF approach. Experienced solvers do it mentally by recognizing letter patterns and common word shapes, but the underlying method is frequency analysis.
Cipher type detection via Index of Coincidence: Frequency analysis tells you more than just the plaintext. It tells you what kind of cipher you are dealing with. The Index of Coincidence (IC) measures how clumped the letter distribution is. English plaintext has an IC around 0.065. A monoalphabetic substitution preserves that value because it just relabels the letters. A polyalphabetic cipher or a transposition cipher flattens the distribution toward random, dropping the IC toward 0.038 (the value for uniform random text over 26 letters).
| Cipher type | Expected IC | What it means |
|---|---|---|
| English plaintext | ~0.065 | Normal clumped distribution |
| Monoalphabetic substitution | ~0.065 | Distribution preserved, letters relabeled |
| Polyalphabetic (Vigenere, etc.) | ~0.038 | Distribution flattened toward random |
| Modern block cipher (AES) | ~0.038 | Indistinguishable from random |
If you compute the IC of an unknown ciphertext and get 0.065, frequency analysis will work. If you get 0.038, you need a different approach. The Letter Frequency Analyzer shows the full distribution chart and the IC value so you can classify the cipher before committing to an attack.
Frequency analysis fails on short ciphertexts. A 30-character sample does not produce a stable frequency distribution. The letter E may not appear at all, or may appear three times by chance. The technique becomes reliable at around 100 characters and highly reliable above 300. For very short texts, pattern-based attacks (word shape matching, dictionary attacks on common short words) work better than raw frequency counting.
Polyalphabetic ciphers, like the Vigenere cipher, defeat naive frequency analysis because the same plaintext letter maps to different ciphertext letters at different positions. The frequency distribution flattens toward uniform. Frequency analysis must be combined with the Kasiski examination and Index of Coincidence to work on polyalphabetic ciphertexts. The standard approach: find the key length using Kasiski or IC, split the ciphertext into groups by position modulo key length, then run frequency analysis on each group independently. Each group is effectively a monoalphabetic Caesar cipher.
Modern block ciphers (AES in any standard mode) produce ciphertext that is computationally indistinguishable from random. Frequency analysis reveals nothing. The distribution is flat by design. This is why classical cryptanalysis, including frequency analysis, is a historical and educational tool rather than a practical attack on modern systems. It still matters because it teaches the core lesson of cryptanalysis: ciphers leak information about the plaintext through statistical structure, and the job of a good cipher is to destroy that structure entirely.
Not reliably. Below about 100 characters, the sample size is too small for the frequency distribution to stabilise. A 30-character ciphertext may not contain a single E in the plaintext, or may contain five by coincidence. The technique becomes consistent above 100 characters and highly reliable above 300. For very short texts, word pattern matching works better than raw frequency counting.
E is the most frequent letter in English, appearing approximately 12.7% of the time in typical text. T is second at 9.1%, A is third at 8.2%. The full frequency order starts ETAOIN SHRDLU, a mnemonic covering the 12 most common letters.
Not directly. The Vigenere cipher uses multiple Caesar shifts in rotation, which flattens the frequency distribution toward uniform. However, once the key length is known (via the Kasiski examination or Index of Coincidence), the ciphertext can be split into groups where each group is a monoalphabetic cipher. Frequency analysis then breaks each group individually.
Al-Kindi, the 9th-century Arab polymath, documented the technique in his manuscript Risalah fi Istikhraj al-Mu'amma around 850 CE. It is considered the oldest known work on cryptanalysis. Al-Kindi applied it to Arabic text, where the frequency distribution differs from English but the principle is identical: count the letters, compare to the known distribution of the language, map accordingly.
ETAOIN SHRDLU is the approximate frequency order of the 12 most common letters in English: E, T, A, O, I, N, S, H, R, D, L, U. It comes from the keyboard layout of Linotype typesetting machines, which arranged keys in frequency order so operators could set type faster. In frequency analysis, mapping ciphertext letters to ETAOIN SHRDLU order is the standard first step.
Letter Frequency Analyzer
Count and analyze letter frequencies in text for cryptogram solving.
Caesar Cipher
Encrypt or decrypt messages by shifting letters through the alphabet.
Substitution Cipher Helper
Tools and utilities for solving substitution cipher puzzles.
Cryptogram Solver
Automated solving of substitution ciphers using frequency analysis and pattern recognition.
Vigenère Cipher
Polyalphabetic substitution cipher using a keyword for enhanced encryption.
The Caesar Cipher: History, Math, and Two Ways to Break It
Julius Caesar shifted letters by 3. Suetonius documented it around 121 CE. Learn the exact math, the ROT13 self-inverse property, and how brute force and frequency analysis break it in seconds.
How the Vigenere Cipher Works, and Why It Was Called Unbreakable
Understand how the Vigenere cipher uses a repeating key to defeat simple frequency analysis, and learn why the Kasiski examination breaks it anyway.
How to Solve a CTF Cryptography Challenge: A Practical Framework
The hardest part of CTF crypto is identifying what you are looking at. Learn the four-step recognition-to-decryption framework for classical, encoding, and substitution cipher challenges.