081 — Historical crawler policy explains an archive blind spot Reviewed September 5, 2026, 20:10 UTC. New evidence Common Crawl's August index contains an independently archived robots.txt for https://paste.ubuntu.com.cn/robots.txt, captured2026-08-13T02:29:19Z. The recovered HTTP200 body includes a CCBot-specific Disallow:/ rule and a final wildcard Disallow:/ group. This gives a concrete, contemporaneous reason that Common Crawl could record the policy without crawling paste bodies. It makes the negative paste-body index result less informative about whether the xinzhai posts existed then. This is an archived host-policy artifact, not an archived xinzhai/xz post. It does not authenticate the paste site's displayed July dates, demonstrate the same alias relationship in August, identify the operator, or show that this policy existed throughout July. No Common Crawl bypass or resubmission was attempted. Archive retrieval and verification One compressed index block covering SURT prefix cn,com,ubuntu,paste)/ was retrieved with HTTP206. It contains exactly one row for that prefix, robots.txt; the next block boundary is beyond the prefix. No numerical paste URL occurs in that exact-host index scope. Crawl:CC-MAIN-2026-34. WARC locator:crawl-data/CC-MAIN-2026-34/segments/1786091385546.31/robotstxt/CC-MAIN-20260813020609-20260813050609-00870.warc.gz Offset874572; compressed length1380. The single corresponding public WARC byte range returnedHTTP206. Recovered WARC target URI, date and record ID match the index. The1862-byte HTTP body hashes to SHA1 Base32 EADKLWRGJMU7ICPNWJBAX75EAHVV2SLA, matching both index digest and WARC payload digest. The local cluster.idx is the same August index used in033; it was not redownloaded. Private raw block, WARC, query/range metadata and digest verification are in081-cc-private. Script081-alias-cc-lookup.py reproduces the index lookup. No other archived content was fetched. Separate public URLscan witness check Two read-only search API calls completed. A query covering domain:paste.ubuntu.org.cn OR domain:paste.ubuntu.com.cn after2026-07-01 returnedHTTP200,total0,has_more=false. An example.com control with the same date condition returned an actual public scan. This is a functioning-search negative for public results under that query, not proof no private/unlisted scans or other witnesses exist. No new scan was submitted. Official search API description:https://urlscan.io/docs/api/ . Raw responses and exact queries preserved in081-private with timestamps/hashes;081-urlscan-check.py reproduces them. Four additional web-index queries for two early IDs with Ubuntu, and exact knowledge/version labels excluding the two paste hostnames, returned no results. Those queries provide no independent witness and are not exhaustive web coverage. Implication for infrastructure and next steps The observed Common Crawl gap has a crawler-policy explanation that an alternative geographic vantage point would not fix. A China-near server may still help access other sites, but buying one would not make these absent archive bodies appear. Existing public copies, repository references, ordinary search indexing or other independent dated artifacts remain the useful routes for this cluster. Do not infer archival absence means the posts were fabricated or newly backdated.