Data sources and credits
Flashbao uses dictionaries, linguistic datasets, language tools and official exam lists. This page shows what we use each source for and gives its licence where available.
We don't use generative AI to write flashcard text. Definitions, example sentences and character notes come from the sources below.
- Definitions, pinyin and measure words
- Word segmentation for learner-submitted text
- Character origins, component roles and book-corpus word frequency
- Word origins, classical citations, related words and the Chengyu listEnglish Wiktionary · Extracted with Wiktextract via kaikki.org
- Character glosses, readings, radicals and character variants
- Stroke-order animation and handwriting quizzes
- Stroke-order character data and visual candidate generation
- Similar-looking charactersUnicode Technical Standard #39 · Single-character Han mappings from confusables.txt
- Character decomposition
- Related words (synonyms, antonyms, hierarchies)
- Additional Mandarin relation lemmas
- Additional Mandarin relation lemmasTUFS Basic Vocabulary Wordnets · Relation-only Mandarin synset memberships
- Word frequency and part of speechSUBTLEX-CH · Cai & Brysbaert (2010), PLoS ONE
- Example sentences
- HSK/YCT/BCT membership, HSK grammar levels and grammar reference dataChinese Testing International · Official HSK grammar cases are retained as seed-time reference data; they are not currently returned by the learner API
- TOCFL deck membershipSteering Committee for the Test of Proficiency-Huayu (SC-TOP) · Source pinyin selects the matching CC-CEDICT reading
Licensing notes
Flashbao reformats, combines and abridges source material. Material from CC-CEDICT, Dong Chinese and Wiktionary is provided under the ShareAlike licences linked above.
A source shown without a licence is credited for transparency; the credit does not imply that the material is openly licensed. Any reuse remains subject to the source's own terms.
Flashbao is independent and is not affiliated with or endorsed by Chinese Testing International, Hanban/CLEC or SC-TOP. Dataset names and trademarks belong to their owners. Found an attribution error? Let us know.