WHAT THE BILLION-PAGE SCAN COVERS Reviewed 2026-09-05 UTC. Common Crawl's official language table reports Chinese (zho) as 4.3829% in CC-MAIN-2026-34, 4.4292% in July (2026-30), and 4.5771% in June (2026-25). The language detector is CLD2; the table uses the primary language and language detection is applied to HTML pages. These are crawl-composition figures, not estimates of what fraction of the Chinese internet is captured. Source: https://commoncrawl.github.io/cc-crawl-statistics/plots/languages The local scanner counted 2,083,525,558 text records. Do not multiply that count by the published language percentage and describe the result as a measured Chinese-page total: the scanner's text-record denominator and the official language table's denominator have not been reconciled. Common Crawl warns that its domain rankings reflect crawling constraints, including robots exclusions and avoidance of excessive server load. Therefore highly ranked real-world domains may be underrepresented. Our own archive work illustrates a narrower practical gap: neither of two dated share-text candidate IDs was present in the inspected August index block, despite successful live listing retrieval. Source: https://commoncrawl.github.io/cc-crawl-statistics/plots/domains.html Inference: a full pass through this large corpus is a valuable rare-string detector on an accessible crawl, but it is not an exhaustive search of Chinese platforms. Missing pages, expiring unlinked notes, login walls, dynamic content, uncertain publication dates and lexical mismatch all limit sensitivity. Chinese-language index coverage and Chinese-lab activity are also different questions: reviewed lab reports describe English/multilingual research over international sources. The current no-confirmed-actor assessment describes the evidence collected so far. It does not justify assigning a probability of nonexistence from the billion-page denominator.