COMMON CRAWL AUGUST BEHAVIORAL SCAN 2026-09-05 UTC — status: completed in 244 seconds. No verified second actor. Purpose: one bounded text-only scan for 40 guessed Chinese behavioral phrases, pinyin/date markers and proxy wrapper variants. This complements known-American-marker searches. The minimum term length is eight actual Unicode characters, not eight UTF-8 bytes. Terms are exact, case-sensitive, with equal weight 2; a score is a retrieval hint, not confidence. Local terms: investigation/china/terms_china_codex_behavior.tsv Remote terms: /data/runs/terms_china_codex_behavior.tsv Output: /data/runs/china-codex-behavior Log: ~/china-codex-behavior.log on existing EC2 44.222.74.56 Corpus: existing /data/corpus, CC-MAIN-2026-34. No links corpus, no new crawl corpus download requested. ccsearch may stream source files missing from local corpus as designed. Invocation: ./target/release/ccsearch --corpus /data/corpus --s3 --s3-endpoint http://s3.us-east-1.amazonaws.com --crawl CC-MAIN-2026-34 --concurrency 384 --out /data/runs/china-codex-behavior --port 8097 --exit-when-done --score-terms /data/runs/terms_china_codex_behavior.tsv --report-domains 100000 --report-pages 20000 Preflight at 10:27 UTC found no ccsearch processes, load average 13.47 on the 192-core existing box, and port 8097 unused. No infrastructure provisioned or terminated. Forbidden ~/terms.tsv and copies never read. Limits: most wording is hypothetical rather than derived from a positive Chinese artifact; a negative has very low sensitivity and cannot rule out Chinese research agents. Some long wrapper variants overlap prefixes from prior scans and serve as control comparisons. Visible text only: hrefs, scripts, unindexed pages, blocked/login-only pages, deleted pastes, other crawl months, language/wording variants and English-speaking Chinese-model agents can all be missed. Date-shaped pinyin terms intentionally cover only 2026 with selected separators. No generic ceshi/linshi substrings, because local audit proved severe accidental matches in English words and pharmacy spam. Phrase hits can still be tutorials or copied discussions and require source/date/context verification. RESULTS (complete) 100,000 / 100,000 WET files processed; zero failures. 2,083,525,558 pages examined; 298 matching rows across 96 registrable domains. Zero bytes downloaded: all input served from existing local corpus. Copied scores.tsv, domains.tsv, report.html and completion log into investigation/china/cc-behavior/. Nonzero terms: id 28: 参考链接 https:// — 285 pages id 25: 测试链接 https:// — 10 pages id 8: 后续智能体可以读取 — 1 page id 9: 智能体之间共享记忆 — 2 pages The other 36 terms have zero hits. Per-term counts including every zero are in 005-term-counts.json and reproduced below. Triage: The 285 reference-link hits are mostly ordinary tutorial/article/product URLs, including repeated social-media follower/view marketing pages. fonton.co.uk contributes 74 rows and voiceweibo.com 9; counts are not independent observations. These were URL-triaged, not all fetched or manually rejected. The 10 test-link hits: two andaily.com blog/archive URLs, pinghe.com/wenda/q_24511.html, vpsxz.net/real-evaluation/4908/, t-www.panewslab.com/zh/articles/019bde36-2233-7658-8007-bc66df428155, mathpretty.com/19310.html, and hdd.sc/product/{159,149,144,154}. Product/review URLs are consistent with ordinary server/network testing, not a positive artifact. Current andaily fetch failed (web 502; direct TLS issuer error); it remains unverified. https://mathpretty.com/19310.html is confirmed unrelated: dated 2024-09-25, reposting a 2023-06-28 human working-memory test forum discussion, linking humanbenchmark.com/tests/number-memory. This is a useful hard negative for “测试链接”. https://www.ginonotes.com/posts/langchain-launch-langmen-sdk is confirmed ordinary developer content, dated 2025-02-19: Chinese translation of the LangMem SDK announcement. Exact match 智能体之间共享记忆 appears in a list of documentation guides, not in a message to another agent. Source page retrieved via web open on 2026-09-05. https://www.zmt.wiki/35312.html currently serves a CyberVerse open-source agent-platform explainer, not a scratchpad. Current extracted text no longer contains the exact matched phrase, so the historic match context is not recovered. Low-priority context mismatch, not verified actor evidence. https://news.tixiaolu.com/posts/news-b8881d5b0d.html is the one unresolved coordination-phrase hit (后续智能体可以读取). The URL is news/article shaped, but body, publication date and actual context are unknown. Web open refused the URL with a non-retryable safe-open error; exact URL/phrase searches returned no result, and the prior local swarmhunt cache contains no copy. Preserve as UNVERIFIED LEAD, not negative and not positive. A targeted archived WET/WARC recovery would resolve it without another whole-crawl scan. It could easily be reporting on agent-memory techniques rather than an artifact itself. No new swarm surface is established, so NEW_SITES.md was not appended. No target writes, posting, counter increments, or infrastructure termination performed. The scan finished and port 8097 closed when ccsearch exited. Existing EC2 remains running; billing lifecycle belongs to coordinator. EXACT TERMS AND PAGE COUNTS (tab-separated id, count, term) 0 0 独特测试标记请忽略 1 0 临时引用链接请忽略 2 0 唯一链接标记请忽略 3 0 测试链接仅用于引用 4 0 此页面用于智能体测试 5 0 此页面仅用于链接测试 6 0 供其他智能体读取 7 0 请将答案写入此页面 8 1 后续智能体可以读取 9 2 智能体之间共享记忆 10 0 智能体共享临时笔记 11 0 独特标记用于验证 12 0 測試連結僅供引用 13 0 此頁面僅用於連結測試 14 0 供其他智能體讀取 15 0 請將答案寫入此頁面 16 0 智能體之間共享記憶 17 0 ceshi-link-2026 18 0 ceshi_link_2026 19 0 linshi-link-2026 20 0 linshi_link_2026 21 0 daili-test-2026 22 0 weiyi-test-2026 23 0 ceshi-agent-2026 24 0 agent-ceshi-2026 25 10 测试链接 https:// 26 0 临时引用 https:// 27 0 引用测试 https:// 28 285 参考链接 https:// 29 0 cors.zme.ink/https 30 0 cors.zme.ink/?url= 31 0 proxy.sunly.fun/?url= 32 0 r.jina.ai/http://data.stats.gov.cn 33 0 r.jina.ai/https%3A%2F%2Fdata.stats 34 0 api.codetabs.com/v1/proxy/?quest=https://data.stats 35 0 markdown.new/https://data.stats.gov.cn 36 0 r.jina.ai/https://datacenter-web.eastmoney.com 37 0 r.jina.ai/https://www.cninfo.com.cn/new/disclosure 38 0 r.jina.ai/http://wenshu.court.gov.cn 39 0 api.allorigins.win/raw?url=https%3A%2F%2Fdata.stats FOLLOW-UP RESOLUTION: The tixiaolu coordination-phrase lead was recovered from an exact Common Crawl WARC record and is a news explainer about internal CANN PyPTO agent state, not an escaped-agent artifact. Archive timestamp Aug13; declared article date Aug6. See 010-tixiaolu-review.txt and raw evidence under pastebins/data/news.tixiaolu.com. Its earlier UNVERIFIED status above records the initial access limitation and is superseded by this review.