963 components build every HSK character
The usual advice is to learn the 214 radicals. Kangxi's 214 are a dictionary index, built to give every character exactly one lookup key, which is a different job from listing what a character is made of. Measured across all 2,630 HSK characters, 963 distinct components appear.
That number sounds worse than it is, because the components are steeply unequal. The 1,421 characters below use nothing outside the 300 most reused ones. Every figure on this page is derived from those files rather than typed in, and it is rebuilt whenever they change.
The components that carry the most weight
| Component | Name | HSK characters containing it | Share |
|---|---|---|---|
| 口 | mouth | 696 | 26% |
| ⺆ | borders radical (form of 冂) | 497 | 19% |
| 丿 | falling stroke | 401 | 15% |
| 人 | person | 354 | 13% |
| 丨 | vertical stroke | 265 | 10% |
| 土 | earth | 239 | 9% |
| 丷 | "eight" component in Chinese characters | 236 | 9% |
| 亠 | lid | 236 | 9% |
| 木 | tree | 235 | 9% |
| 日 | sun, day | 222 | 8% |
| 勹 | wrap | 211 | 8% |
| 扌 | hand | 170 | 6% |
The head of the list is not made of meanings. 口 is a mouth in 吃 and carries nothing at all in 兄. Strokes and enclosures rank high because they are shapes, not ideas. This is the honest version of the mnemonic advice: components make characters easier to remember and write, and only sometimes easier to guess.
What the top of the list buys you
A character counts here only when every component under it, all the way down, is inside the set. It is a strict test, which is why the numbers climb slowly and then quickly.
| Learn the N most reused components | HSK characters built entirely from them | Share of the list |
|---|---|---|
| 25 | 156 | 6% |
| 50 | 284 | 11% |
| 100 | 538 | 20% |
| 200 | 1,017 | 39% |
| 300 | 1,421 | 54% |
| 500 | 1,978 | 75% |
300 components is a fortnight of work and it fully accounts for 54% of the HSK characters. Going from 300 to 500 adds 557 more. Past that the return collapses, which is the argument for stopping and learning the rest of the tail inside the characters that use them.
The other half: what a character builds
Components are what goes into a character. Words are what comes out. Over the 13,563 words in this dictionary, written with 3,729 distinct characters:
The median character appears in 3 words and 951 characters, 26% of them, appear in exactly one. So the distribution has a very short head and a very long tail, and the head is worth learning as vocabulary rather than as characters.
| Know the N most productive characters | Words you can spell entirely | Share of the dictionary |
|---|---|---|
| 100 | 744 | 5% |
| 300 | 2,623 | 19% |
| 500 | 4,313 | 32% |
| 1,000 | 7,554 | 56% |
| 1,500 | 9,620 | 71% |
Spelling a word is not knowing it. 好, 意 and 思 do not give you 不好意思, and that is the point of the table rather than a caveat on it: characters put words within reach and do not deliver them. The coverage version of the same question, measured on real prose, is on how many characters you need.
What this measurement does not tell you
- A decomposition is an analysis, not a fact. Reasonable sources split the same character differently, and how deep to recurse is a choice. A different decomposition would move 963 by some tens.
- 81 HSK characters carry no decomposition here. They are counted in the 2,630 and can never be counted as built from a component set, so the share column is if anything pessimistic.
- A component is not a meaning. Semantic components hint at sense, phonetic ones hint at sound, and plenty are neither. Nothing in the counts above distinguishes the three.
- The word figures are this dictionary's, not the language's. 13,563 entries is a working vocabulary list, not a full lexicon, so treat the shares as shares of it.
- 214 is not wrong, it is a different question. If you want to look a character up on paper or understand how a dictionary is ordered, the Kangxi radicals are the right thing to learn. They were never an inventory of parts.
Questions about components and radicals
How many radicals do you need to learn for Chinese?
Fewer than the inventory and more than 214. 963 distinct components appear across the 2,630 HSK characters, but they are steeply unequal: the 100 most reused fully build 538 characters and the 300 most reused fully build 1,421, which is 54% of the list.
Are there 214 radicals in Chinese?
There are 214 Kangxi radicals, and that is an indexing scheme for finding a character in a paper dictionary. It assigns each character exactly one radical, which is a different job from listing what the character is made of. On the decomposition used here, 963 components actually appear, 939 of them as a direct part of an HSK character.
Which component appears in the most Chinese characters?
口, in 696 of the 2,630 HSK characters. Then ⺆ (497), 丿 (401), 人 (354).
Which Chinese character appears in the most words?
不, in 216 different words in this dictionary of 13,563. The median character appears in just 3, and 951 characters appear in exactly one. Productivity is distributed far more unevenly than frequency is.
Is learning components worth the time?
For the top of the list, yes, because the payoff is measurable: 538 HSK characters use nothing outside the 100 most reused components. For the tail it is not. Roughly half of the 963 appear in a handful of characters each, and meeting those in the characters themselves costs less than studying them separately.
Where does this data come from?
The character decomposition graph and HSK character list this site serves, the same files behind the dictionary and the stroke animations. Every figure is recompiled from those files whenever they change, so nothing here is a number somebody typed in.
300 components account for 54% of the HSK characters, and the other 663 are best met inside the characters that use them. GraphChinese orders them that way, and brings each one back the day before you forget it. The whole course, HSK 0 through HSK 6, is $49 once for lifetime access.
Start at your true level