Chinese lab agent capabilities: primary-source hypotheses, not swarm evidence Reviewed 2026-09-05 UTC ASSESSMENT There is strong primary-source evidence that several Chinese labs train agents with live web-search tools and code-execution environments. There is no evidence in these sources that those agents used anonymous public pages for shared memory, created citation URLs on third-party writable surfaces, or left the artifacts sought in this investigation. A model's tool competence, its particular execution environment, and an observed public artifact are three separate propositions. Four examples below are sufficient to refine the search without relying on product marketing. Each uses a versioned technical report predating the March–August 2026 investigation window. “High confidence” means confidence in what the report describes, not independent replication or lab attribution. All summaries are paraphrases; no direct quotations are necessary. 1. ALIBABA / TONGYI DEEPRESEARCH — documented retrieval plus sandbox computation Source: Tongyi DeepResearch Technical Report, 28 October 2025, sections 3.4.3 and Appendix D. https://arxiv.org/html/2510.24701v1#S3.SS4.SSS3 https://arxiv.org/html/2510.24701v1#A4 Official repository: https://github.com/Alibaba-NLP/DeepResearch Observed: the report specifies Search, Visit, Python Interpreter, Google Scholar and File Parser. Search uses Google; Visit parses pages with Jina before a separate model extracts goal-relevant information. Python executes inside a sandbox. The paper also describes simulated and real environments rather than one uniform execution setup. Its evaluation spans BrowseComp and BrowseComp-ZH alongside academic and general information-seeking tasks. Inference: Google, Jina, English queries and international information sources are plausible search directions even for this Chinese lab. Language-only searches would miss plausible activity. The specified retrieval tools do not provide an explicit form-submission or generic HTTP-method interface. Unknown: sandbox outbound-network policy, allowable URL schemes/parameters, precise production tool implementation, public-write permissions, egress provider, and whether any deployment used third-party pages as memory. Python capability alone does not resolve these unknowns. Confidence: HIGH for published tool inventory; LOW for any public-write hypothesis. 2. MOONSHOT / KIMI K2 — simulation must not be mistaken for public activity Source: Kimi K2: Open Agentic Intelligence, 28 July 2025, sections 3.1.1 and 3.2.1; Appendix C. https://arxiv.org/html/2507.20534v1#S3.SS1.SSS1 https://arxiv.org/html/2507.20534v1#S3.SS2.SSS1 Observed: the synthesis pipeline draws specifications from over 3,000 real MCP tools and over 20,000 synthetic tools. A tool simulator maintains state and generates feedback; actual execution sandboxes supplement simulation for coding and software engineering. The report describes Kubernetes infrastructure supporting more than 10,000 concurrent sandboxes. Tasks include GitHub issue resolution and English/Chinese API-use evaluation. Inference: “thousands of agents” in a training report is not evidence of thousands of agents writing to the open internet. Simulated state changes can resemble real service operations in a transcript. Genuine executable coding environments make terminal-related capability plausible but do not identify their network permissions. Unknown: unrestricted outbound HTTP from those sandboxes, live service credentials, public browsing during these particular rollouts, deployment location, and any third-party scratch-page use. Confidence: HIGH for the simulation/real-execution distinction; LOW for externally observable swarm claims. 3. DEEPSEEK V3.2 — real search APIs and multilingual synthetic research tasks Source: DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models, 2 December 2025, section 3.2.3 and sections 4.1/4.4. https://arxiv.org/html/2512.02556v1#S3.SS2.SSS3 https://arxiv.org/html/2512.02556v1#S4.SS4 Observed: the report distinguishes real search, software-engineering and Jupyter environments from synthesized general-agent environments. It lists 50,275 search tasks, 24,667 coding tasks and 5,908 interpreter tasks. Search questions are synthesized around long-tail entities using multiple agents; the resulting data covers multiple languages and domains. The coding environment-setup agent performs package installation, dependency resolution and tests. Search evaluation uses a commercial search API; context management may summarize or discard earlier tool history. Inference: broad international facts, rare entities and repeated self-verification are better task-domain hypotheses than a narrow focus on Chinese financial endpoints. Search and coding capabilities must not be assumed to coexist in every rollout. Package installation establishes an environment-setup operation, not unrestricted network access for every subsequent model action. Unknown: actual HTTP tool schemas, public-write routes, sandbox network policy, egress IPs, and any external memory mechanism. Confidence: HIGH for documented live search; LOW for public-write attribution. 4. ZHIPU / Z.AI GLM-4.5 — search RL and isolated software environments Source: GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models, 8 August 2025, sections 3.3, 3.4 and 3.5. https://arxiv.org/html/2508.06471v1#S3.SS3 https://arxiv.org/html/2508.06471v1#S3.SS5 Observed: web-search training uses difficult questions assembled across multiple webpages, including knowledge-graph construction and human-assisted extraction/obfuscation. Software tasks derive from GitHub issues and pull requests and execute with tests in isolated sandboxes. The infrastructure provides concurrent Docker task environments and asynchronous agent rollouts. General function-calling training also uses MCP-based synthesized tasks and runnable environments. Inference: multi-hop fact retrieval and software debugging are documented task families; this motivates looking for coherent research sequences rather than the literal Chinese word for “agent.” A unified HTTP interface in the RL infrastructure is an internal integration mechanism, not evidence that agents have an arbitrary outbound HTTP tool. Unknown: exact search adapter, terminal permissions, unrestricted outbound connectivity, region/provider of execution, and whether persistent public pages were ever used. Confidence: HIGH for training architecture; LOW for a third-party-public-surface hypothesis. HOW THIS CHANGES THE HUNT — ANALYST INFERENCES A. Separate three environment hypotheses in every lead: (i) search-result API only; (ii) arbitrary URL retrieval/visit, possibly mediated by a reader service; (iii) terminal, code interpreter, browser interaction or generic HTTP tools. The second could in principle trigger a poorly designed GET-writing endpoint, but only if its adapter passes such URLs through. The third expands possible request methods only if networking and tool policies permit them. Neither is permission evidence or proof that writes happened. Our investigation remains read-only. B. Search for multilingual and English task traces as well as Chinese traces. Tongyi's explicit Google/Jina stack and DeepSeek's multilingual task construction undermine a China-domains-only strategy. These facts do not establish any model's runtime geography. C. Prioritize artifacts that resolve the environment uncertainty: dated public logs showing an actual sequence, distinctive cross-site identifiers, a reproducible page history, or a sufficiently specific tool/request signature. A model name, generic test string, cloud ASN, benchmark name or framework token is not enough. D. Keep training-simulation artifacts in a separate category. An example that calls a synthetic “publish” tool is evidence of a simulated API interaction until execution against a real public service is independently demonstrated. E. Long-context handling offers a reason to investigate memory-related behavior, but the reviewed reports already describe internal solutions such as summarization or history discarding. External scratch memory is a hypothesis to test, not a necessity implied by context limits. SCOPE AND REMAINING QUESTIONS This is a selected four-example technical review, not an exhaustive survey or a comparison of the newest releases. ByteDance Seed2.0 and MiniMax official repositories were located during discovery, but their detailed environment policies were not verified sufficiently for a fifth or sixth example. They remain future documentation work, not negative evidence. No reviewed source establishes a Chinese agent swarm on public writable surfaces. None justifies inferring runtime IP ranges, mainland location, corporate ownership of an unknown artifact, or the absence of relevant activity. The next decisive evidence must come from the artifacts themselves.