MiroFlow: recovered public hierarchical-agent execution trace Reviewed 2026-09-06 UTC Finding The broken README link did not mean the public trace collection was gone. Current benchmark documentation points to a publicly accessible, non-gated Hugging Face archive. One sampled record contains a main agent, four worker session histories and structured tool calls. This is an actual published run artifact, stronger than framework descriptions, though not independently authenticated model-provider telemetry or evidence of an escaped swarm. Primary provenance and download Official repository: https://github.com/MiroMindAI/MiroFlow Reviewed Git tree: fca7b4a7fda30e5a17b5c7b73d780738a925869f Relocated instructions: https://github.com/MiroMindAI/MiroFlow/blob/fca7b4a7fda30e5a17b5c7b73d780738a925869f/docs/mkdocs/docs/gaia_validation_claude37sonnet.md Public archive: https://huggingface.co/datasets/miromind-ai/MiroFlow-Benchmarks/resolve/main/gaia_validation_miroflow_trace_public_20250825.zip Size: 27373060 bytes (27.37 MB) SHA256: a86e519e5c604f277ec0a010284a2426c702796e4eeedba57d84f8d61597bc55 Both size and SHA256 were checked against Hugging Face LFS metadata. Dataset API explicitly reports gated=false. ZIP members are password-protected, but the same official download instructions publish the passcode. It was used as provided. No gated MiroVerse files or terms acceptance were involved. Archive inventory: 168 members, including one directory, 165 task JSONs, one JSONL file and one TXT file. Only one task JSON was decoded and behaviorally inspected in this bounded pass. Inventory counts do not authenticate the claimed benchmark score or prove completeness against the original benchmark. Sample Member: run_public/task_e2d69698-bc99-4e85-9880-67eaccd66e6c_attempt_1_2025-08-16T11-52-45-514409Z.json It concerns finding a Survivor winner and records an August 16, 2025 run. Fields include main_agent_message_history and sub_agent_message_history_sessions. Four workers are labeled agent-worker_1 through agent-worker_4, with 23, 47, 17 and 45 stored messages respectively. Distinct histories and delegation calls establish a recorded hierarchy, not that all workers ran simultaneously. The sample itself marks its answer correct; that judgment was not independently evaluated here. Observed assistant tool calls, excluding instructions and embedded examples 61 total: 35 tool-searching/google_search 13 tool-searching/scrape_website 4 agent-worker/execute_subtask 4 tool-searching/wiki_get_page_content 2 tool-code/run_python_code 1 tool-reasoning/reasoning 1 tool-code/create_sandbox 1 tool-code/download_file_from_sandbox_to_local The main agent delegates research subtasks, while worker histories contain searches, page retrieval and local data preparation. The sample's code writes Survivor data files under /home/user/ in a sandbox, then downloads a result file locally. Those code strings were inspected only, never executed. External URL operations Explicit scrape arguments include Gold Derby, IMDb, Survivor Fandom, Reddit, a Wikipedia mirror at wikipedia.nucleos.com, a Weebly page and Famous Birthdays. The URLs were read as trace data, not visited. No paste upload, wiki edit, anonymous public storage write or check-in request was observed in this sample. The Wikipedia-named tool is a page-content read, not evidence of wiki editing. This one sample cannot establish absence elsewhere in the 165-task collection. Potential corpus discriminators The structured strings agent-worker/execute_subtask, tool-searching/scrape_website and tool-code/download_file_from_sandbox_to_local are better candidates than a generic phrase such as agent swarm. They still may be shared by copies of the framework, documentation or training material; a literal match alone would not attribute a public artifact to MiroMind. Attribution and scope This is publisher-linked MiroMind research evidence. The documentation labels this benchmark run as using Claude 3.7 Sonnet, so Chinese developer/lab interest must be separated from base-model origin. This pass does not establish team location, execution egress or a connection to xinzhai/xz. Its relevance is that published hierarchical research traces provide concrete behavior to compare against unexplained public-web artifacts. Private artifacts: 156-private contains Git tree/docs, Hugging Face metadata, hash-verified ZIP, member inventory, sample and sample provenance, actual-calls.json and SHA256 inventory. An earlier calls.json also includes five tool-formatting examples and must not be used for actual-call totals; actual-calls.json filters to assistant-role messages and is the authoritative count for this report. No posting, accounts, code execution, monitoring pings or site edits performed.