028 — REVIEWED FINDINGS: local Ubuntu paste cluster structure Review date: 2026-09-05. Fixed download-metadata snapshot captured 11:32:14 UTC. Finding The xz_knowledge_p1 posts form a strong automated posting candidate: 233 distinct bodies, overwhelmingly five-minute cadence. A nearby, temporally interleaved xinzhai family contains nine increasingly large versioned, multipart payloads that decode into Fernet-layout-compatible tokens. This is concrete evidence of patterned public storage behavior, plausibly an encrypted storage client or prototype. It does not establish an autonomous research agent, multiple collaborating agents, laboratory affiliation, Chinese authorship, or escape from an evaluation environment. The relationship between the two naming families is suggestive, not authenticated. Coverage and reproducibility This analysis made ZERO origin requests and did not modify, stop, or restart the ongoing numeric survey. It examined only completed metadata records and hash-verified local raw HTML. The metadata snapshot has 829 records: 828 valid paste bodies and one excluded position. All 828 parse successfully. The contiguous downloaded prefix is IDs 4548081–4548899, 819 bodies; nine additional valid sparse anchors lie beyond it. Displayed dates across the contiguous prefix run 2026-03-17 03:51 to 2026-07-12 09:31. The sparse anchors extend displayed date coverage to September 5, but do NOT supply continuous July–September coverage. Source URL template: https://paste.ubuntu.org.cn/. Raw private directory: pastebins/data/paste.ubuntu.org.cn/china-codex-survey/. Exact metadata snapshot: 028-ubuntu-metadata-snapshot.jsonl; SHA256 b9402942a3c9797afcda727e726c7220ebfcca7f17919f7aed9d2fbe49dfcf65. Every reviewed page's URL, fetch time and raw hash are retained there. Known sensitive excluded ID remains excluded; no secret values or opaque payloads are reproduced here. Reproduce the fixed snapshot: python3 investigation/china/ubuntu_cluster_snapshot.py investigation/china/028-ubuntu-metadata-snapshot.jsonl python3 investigation/china/ubuntu_neighbor_shapes.py python3 investigation/china/ubuntu_neighbor_nested.py The first script measures exact authored textarea text, preserving leading/trailing spaces in the space-to-plus decoding hypothesis. It writes internal structured metadata, not payloads. The latter scripts perform bounded static parsing/base64 operations only. No pasted code executes and no embedded URLs are visited. Rerunning the first command refreshes the analysis timestamp, not the fixed input coverage. Author and payload distribution 828 bodies have 271 distinct displayed author strings; these are user-supplied labels, not verified accounts. Largest groups: xz_knowledge_p1 233 (28.1%), Anonymous 209 (25.2%), kyphen 20, test 16, yexiaoyu 9. Grouping xinzhai and numbered/versioned variants yields 79 posts (9.5%). One additional xz_improvement_plan_p1 post appears; no other xz_ variants occur in this snapshot. 315/828 bodies meet a deliberately broad base64 alphabet/length screen. After canonical decoding and entropy/non-UTF8 checks, 237 meet the operational opaque-body classifier: all 233 xz_knowledge_p1, the one xz_improvement_plan_p1, and three Anonymous posts. This classifier measures shape, not encryption or purpose. The remaining 513 fail the initial alphabet screen; 78 screen-positive bodies are not opaque by these criteria. An initial run stripped edge spaces and falsely rejected four xz records; preserving exact edge spaces resolves all four. Reported counts here use the corrected method. 233/233 xz_knowledge_p1 bodies are unique hashes, one line, tagged python. Lengths: 178 at 124 characters, 52 at 424, one at 488, two at 500. Displayed author dates span July 10 22:24 to July 16 09:34 including sparse anchors. The 230 contiguous-prefix posts produce 229 successive displayed-time intervals: 184 exactly five minutes (80.3%); median five, 14 zero-minute intervals, one 936-minute gap, and the rest 1–20 minutes. Dates never reverse within this group. Minute-resolution timestamps and missing later coverage limit interpretation. All 234 xz_ family bodies become canonical standard base64 when literal spaces are restored to plus signs. None has a literal plus. Decoded sizes are 93 bytes (178 posts), 316 (52), 364 (one), 374 (two), and 646 (the improvement-plan label). All 234 lack the Fernet version byte 0x80 and fail the tested Fernet token layout. They cannot be identified as that format using this layer; absence of this format is not evidence against some other encryption or encoding. No timestamp was guessed from arbitrary bytes. The form-urlencoded plus-to-space hypothesis is plausible but not proven against original submitted bytes. The three earlier Anonymous opaque bodies (4548186, 4548187, 4548193; June 4–6) differ materially: 53–79 lines, 4,096–6,122 characters, preserved literal plus signs, and zero spaces. Thus base64 sharing is broader than this handle, while its short, single-line, plus-damaged five-minute series is distinctive in the observed corpus. Before the first xz post: 3 opaque bodies among 483 prefix entries. From first xz through prefix end: 231 among 336. Neither denominator is a random sample of all users or all time. Neighboring multipart family Sources begin https://paste.ubuntu.org.cn/4548523 and end https://paste.ubuntu.org.cn/4548609. Three initial xinzhai pastes, July 10 21:26–21:27, are ordinary hello/hello-xinzhai print calls. Static AST parsing recognizes print; it does not execute them. They establish a small publishing test immediately before the multipart sequence, not an agent identity. The next 76 posts have numbered part labels, all tagged python. 67 parts are exactly 30,000 characters, with nine shorter final parts. Group counts and assembled binary sizes: xinzhai: 4 parts, 64,985 bytes, displayed start 21:27 xinzhai_v5.2: 4 parts, 65,545 bytes, 21:32 xinzhai_v52: 6 parts, 99,737 bytes, 21:53 xinzhai_v60: 7 parts, 113,609 bytes, 22:06 xinzhai_v61: 8 parts, 125,353 bytes, 22:12 xinzhai_v70: 9 parts, 151,561 bytes, 22:21 xinzhai_v71: 11 parts, 180,841 bytes, 22:29 xinzhai_v72: 13 parts, 209,961 bytes, 22:35 xinzhai_v73: 14 parts, 231,225 bytes, 22:43 All dates above are July 10, 2026, literal site display. Each group consists of ordinary base64 chunks. Decoding each and concatenating in numeric part order yields a canonical URL-safe base64 token. Decoding that layer yields a binary object with byte 0x80, an eight-byte plausible timestamp, and ciphertext length compatible with Fernet's block layout. All nine match this structural test; their zlib compression ratios are about 1.0003–1.0004. This is stronger format evidence than entropy alone, but no MAC authentication or decryption was performed. No key or password search was attempted. The payload's actual content and producing software remain unknown. The nine embedded, UNAUTHENTICATED timestamps run July 10 13:27:33–14:43:04 UTC. Adding eight hours aligns them with the displayed start times within roughly a minute. This independently supports the earlier hypothesis of UTC+8 display despite the RSS +0100 label, but tokens can be forged or generated before posting; clock provenance and timezone are still unverified. Do not normalize the full corpus to UTC on this evidence alone. xz_knowledge_p1 begins immediately after v70, and its posts interleave within the v71/v72 part batches. Shared python tag, related-looking prefixes, close timing and interleaving support a possible common publishing workflow. The incompatible inner formats prevent a stronger claim of identical software. An encrypted snapshot uploader plus a periodic small-state writer is a plausible behavioral interpretation, alongside unrelated users, testing, backup/storage or other automation. Version names and knowledge/improvement words are labels, not demonstrated capabilities. Next useful local work Once the ongoing survey advances, rerun against a newly preserved snapshot to determine sustained cadence, size transitions, later xinzhai versions, related labels, and any plainly readable explanatory posts. Seek externally documented client code only through normal public search; no decryption guesses are needed. Report 024's six later sources extend observed xz dates to July 20, but those separate forensic files were deliberately not mixed into this fixed survey denominator. Raw payloads and decoded token bytes remain private. Primary format reference, checked 2026-09-05 11:41 UTC The Fernet specification defines a URL-safe base64 token containing version 0x80, a 64-bit big-endian timestamp, a 16-byte IV, block-aligned ciphertext and a 32-byte HMAC. Its verification procedure also requires recomputing and comparing the keyed HMAC. Our layout test does not perform that authentication and therefore cannot certify a valid token or trustworthy timestamp. https://github.com/fernet/spec/blob/master/Spec.md Archived primary source: 034-sources/0.raw, SHA256 87a714a18a9e23ca69b80a45cbccb792c8fac9d5719417bcdf094db153af2d6c.