GraphChinese

Credits

Open work this course is built on, and the people who made it.

Pronunciation audio

The tone and pinyin listening exercises use native-speaker recordings from the audio-cmn open corpus, part of the Shtooka / SWAC project. Syllables recorded by Chen Wang; word recordings by Yue Tan. Compiled by Hugo Lopez.

Licensed CC BY-SA 3.0. These recordings are served unmodified.

A few words absent from that corpus — 你好, 很好, 吃饭 and 中国 — use recordings by Luilui6666 from Lingua Libre, via Wikimedia Commons. Licensed CC BY-SA 4.0. Modified: converted from WAV to MP3 and trimmed of leading and trailing silence.

Character stroke data

Stroke-order animations use Hanzi Writer and its Make Me a Hanzi data (LGPL / Arphic Public Licence).

Character decomposition, graph & example sentences

The character-decomposition trees, the character↔word co-occurrence graph, and part of the example-sentence corpus in our Dictionary are derived from HanziGraph by Matt Reichhoff (MIT licence), which in turn compiles several open datasets:

例句 · Real sentences flashcard decks

The six 例句 decks are built from the Tatoeba Project, licensed CC BY 2.0 FR. The Chinese sentences and their English translations were written by Tatoeba's contributors; we select and grade them so every word in a deck is one this course has already taught, and we add nothing to the sentences themselves.

Dictionary definitions

English definitions and pinyin come from CC-CEDICT, published by MDBG and licensed CC BY-SA 4.0 (referenced work: CEDICT, Copyright © 1997, 1998 Paul Andrew Denisowski). Our Dictionary carries all 124,727 entries. Modified: readings converted from numbered syllables to tone marks, senses split apart, classifier fields extracted, and a small number of readings corrected where the source is wrong. We do not redistribute HanziGraph's definition files.

Word frequency

Search results are ordered using a word-frequency list from the University of Leeds Centre for Translation Studies internet-zh corpus (281 million tokens).

Idiom origins

The 成语 explanations and classical sources shown on idiom entries come from an open Chinese idiom dataset compiled from public-domain reference works.

← Back to GraphChinese