Home›Typing guide›Wubi
Wubi short codes and phrase input — where the real speed comes from
A beginner types every character with its full four-key code and wonders why Wubi does not feel fast. It is not supposed to be fast that way. The full code exists to teach you the structure; the speed comes from two mechanisms layered on top of it.
Three kinds of code, and when each applies
| Kind | Length | Use |
|---|---|---|
| Full code (全码) | Up to 4 keys | Learning the structure; anything not covered by a short code |
| First-level short code (一级简码) | 1 key | The 25 most frequent characters, one per key |
| Second and third-level short codes | 2–3 keys | A few hundred more frequent characters |
| Phrase codes (词组) | 4 keys for the whole phrase | The actual source of speed |
The critical distinction: the first-level short codes are assigned by frequency, not by structure. The character on G is not there because it contains the radical 王 — it is there because it is the most common character assigned to that slot. So you cannot derive short codes from the radical chart; they are a separate small list to memorise.
This is also why the radical drills use full codes while real typing uses short ones. Beginners who learn only the short codes never internalise the structure; beginners who never learn the short codes type at half speed forever.
Learn both, and know which you are looking at
A common early confusion: you look up a character, find code RG, type it, and get nothing — because your input method is expecting the full code, or vice versa. When a source lists "the Wubi code" without saying which kind, it is usually the full code.
Practical rule for learning:
- Learn the full code of a character first. It teaches the split.
- Then check whether a short code exists, and use it from then on.
- Do not bother memorising second and third-level short codes deliberately — they appear often enough in phrase input that you absorb them.
Phrase input: the actual speed mechanism
Wubi's real advantage is that you can type a whole phrase without selecting candidates. The code for a multi-character phrase is built from components of its characters:
- Two characters: first two codes of each → 4 keys
- Three characters: first code of the first two characters + first two codes of the last → 4 keys
- Four or more characters: first code of each of the first three + first code of the last → 4 keys
So a four-character idiom costs four keystrokes, exactly like a single character. At that rate a four-character phrase takes the same effort as one character, which is why experienced Wubi typists reach speeds that pinyin typists cannot match on the same text.
Two consequences worth internalising:
- Type in words, not characters. Never type one character, confirm, then the next. Type the phrase as a unit and let it come out whole. This is the single biggest behavioural difference between a slow Wubi typist and a fast one.
- Do not watch the candidate window. If your phrase code is right, the phrase appears. Watching the window is a pinyin habit and it breaks the rhythm.
Why Wubi beats pinyin on candidate selection
Pinyin input requires you to choose from homophones. Even with good prediction, a single syllable like *shi* has dozens of candidates, and the IME's guess is wrong often enough that you are reading a list several times per sentence.
Wubi codes are near-unique, so the candidate window is usually empty or has one entry. You are not selecting; you are typing. That is the throughput difference, and it is why Wubi survives in professional transcription despite pinyin's overwhelming popularity.
The flip side: Wubi requires you to know how to write the character, and gives you no help with pronunciation. For a learner, or for someone who encounters text they cannot read, that is a real cost.
A realistic progression
- Weeks 1–4: Full codes, single characters, accuracy only. Speed will be poor and that is expected.
- Months 2–3: First-level short codes become automatic; begin typing two-character words as units.
- Months 4–6: Phrase input becomes the default. This is where speed crosses pinyin-level throughput.
- After that: Gains come from vocabulary of phrases rather than from technique.
Where the code data comes from
The Wubi codes used in this site's lessons come from open-source code tables: the 86 version (Apache-2.0), the 98 version (Unlicense) and the New Century version (MIT). The codes themselves are unchanged from those sources; only the subset of characters and phrases is selected.
Two practical notes that follow from this:
- Short codes are version-specific. If you switch versions, the short code list changes and must be relearned.
- Phrase tables vary between input methods. A phrase your IME does not have will fall back to character-by-character input, which feels like a mysterious slowdown. Adding it to your personal dictionary fixes it permanently.
Run the phrase lessons rather than the single-character ones. Type the code without waiting for the candidate window, and watch your characters-per-minute jump.
Open the practice panel