A Family Reunited Across Millennia
The Sino-Tibetan language family — encompassing Chinese, Tibetan, Burmese, and hundreds of smaller languages — is one of the largest on Earth by speaker count. Yet compared to Indo-European, it remains poorly understood at the proto-language level, partly because Chinese's long written history tempts scholars to rely on characters rather than sounds, and partly because the family's internal classification is still debated.
The most powerful evidence linking Chinese to its relatives is vocabulary cognates: words descending independently from a common ancestral word, with sound changes that follow predictable rules.
The Number Five: *ŋā Across Three Branches
The Proto-Sino-Tibetan (PST) root for 'five' is reconstructed as *ŋā. Here is how it surfaces across branches:
- Chinese: Old Chinese *ŋˤaʔ (pharyngealized, with glottal suffix) → Middle Chinese *ŋuoX → Mandarin wǔ 五. The initial *ŋ- was simply lost: the 疑 initial merged into Mandarin's zero initial, and the w of wǔ is only pinyin's way of writing a syllable that begins with u. The same loss gives 銀 yín and 我 wǒ, whichever way the vowel rounds.
- Tibetan: lnga (ལྔ). The l- is a prefix, not part of the root; the core is ŋa, directly matching the PST root.
- Burmese: ngā (ငါး). The initial ŋ- is preserved without the prefix complication, and the vowel directly reflects *ā.
The Number Four: *li(j) Across Three Branches
The root for 'four' reconstructed as *li(j) shows an equally striking convergence:
- Chinese: Old Chinese *s.li[j]-s → Middle Chinese *sijH → Mandarin sì 四. The *s- initial replaced *l- through a prefix interaction (compare the *s.rǝn pattern for 山).
- Tibetan: bzhi (བཞི།). Again, the b- is a prefix; the zh- of the root is Tibetan's own palatalisation of the PST lateral *l- before the front vowel of *li(j).
- Burmese: le (လေး). The l- directly preserves the PST *l- initial.
Two More Cognate Pairs
Eye / 目: PST *myak → Old Chinese *C.m(r)[u]k → Mandarin mù; Tibetan mig (མིག); Burmese myak. The core consonant *m- is rock-stable across the family.
Name / 名: PST *r-miŋ → Old Chinese *meŋ → Mandarin míng 名; Tibetan ming (མིང་); Burmese a-maññ (အမည်, pronounced /ʔəmjì/). A nasal final is shared across all three branches — Burmese has a palatal -ññ where Chinese and Tibetan have -ŋ, a regular correspondence.
The Critical Methodological Caveat
Identifying cognates is harder in Sino-Tibetan than in Indo-European for three reasons. First, Chinese's monosyllabic surface structure and the logographic writing system hide older morphology. Second, Chinese and Tibetan were geographically adjacent for millennia and exchanged loanwords freely — a resemblance might be borrowing, not inheritance. Third, the family's internal tree is contested: is Tibetan closer to Chinese, or are they both branches of a polytomy? Until the tree is resolved, directional sound-change rules are hard to establish. Sagart's controversial Sino-Austronesian hypothesis — that Chinese and Austronesian share an ancestor — shows how much remains open. It is easily confused with Austro-Tai, which is a different claim by a different scholar: Benedict's 1942 proposal joining Austronesian to Kra-Dai, with Chinese not in it at all.
Despite these caveats, the numerals are the most secure cognates: they resist borrowing (you rarely borrow counting words) and the sound correspondences are regular, making 'five' the strongest single thread stitching the family together across 5,000 years.