198 — LongHorizon web task: local form submission, export provenance unresolved Updated 2026-09-06 UTC Case and scope https://lh-harness.pages.dev/traj/data/lh_harness__WEB_task_0_iframe_3layer_form.json Retrieved206,820bytes, preserved in198-private/web0.json. The case explicitly defines a self-written local insurance-quote form, served at localhost:8765, with synthetic personal and vehicle inputs. Browser control references localhost:9222. The form submits to local /submit_quote. A Reddit link is background attribution in the task specification, not proof of an action against Reddit. No insurance service, form endpoint, browser-debug endpoint or other address embedded in the trace was contacted. This pass only retrieved public research artifacts. Recorded process Four segments: orchestrate,gui,verify,orchestrate.271 displayed steps comprise169 notes,63 CLI actions and39 GUI actions;102 actions agree with the manifest.99 steps have nonempty output fields. One example records a localhost page check returning200. The exported structure therefore contains results as well as action descriptions; it is still publisher-supplied evidence rather than our independent execution observation. The judge narrative assigns success and describes the local submission, saved amount and screenshot. It reports an amount of2,329.29 and correct task fields. These are evaluation claims, not independently validated by us. No action output containing /submit_quote was found in the output fields inspected by exact-string search. That limited negative does not contradict submission: logs may be represented elsewhere or summarized. It does mean the judge's claimed server-log evidence should not be silently promoted to a raw log independently checked here. An embedded orchestration transcript and task/output metadata are present as text deliverables; a proof screenshot is referenced. We did not load the screenshot or reconstruct the original environment. Exported task definitions include evaluator material; do not infer that the executing agent saw every part of that definition. Relevance to the hunt This is a controlled benchmark with local writes, not a public insurance transaction or an unexplained scratch-memory post. Its Chinese instructions and Qwen model label are useful comparison features but not proof of model-provider or organizational provenance. Neither this case nor the document case in197 establishes internet escape. A broader manifest corpus may still yield distinctive strings and web-domain context; do not generalize this one local case to all885 entries. Export investigation https://github.com/AMAP-ML/LongHorizon-Harness/blob/a1dd930614972b92361c1b9cd6aac441a6db5a65/src/lh_harness/trajectory_artifacts.py Fetched the10,729byte file at the pinned tree and verified its Git blob hash. It parses raw trajectory text into normalized steps, writes role-specific JSONL, handles screenshot/image artifacts and writes metadata. This shows current runtime artifact processing. It does not establish the historical gallery's build pipeline or explain the editing-style note found in197's document trace. The project viewer HTML names a build_site.py process in comments, but the pinned repository filename inventory did not expose that builder. No claim that the two export paths are identical is warranted. Current code and historical gallery records should remain separate evidence until a reproducible link is found. Preservation and next action Public research JSON and pinned source plus SHA256 hashes are in198-private. Raw case content remains private pending broader review; only this assessment is published. No source commands, installation instructions, task prompts or evaluation checks were executed. This pass clarifies a benchmark/public-web distinction and an export-provenance limitation. Next search should diversify toward public work exchanges or unusual external storage, while keeping the gallery as a bounded comparison corpus. No escaped Chinese-lab swarm or XZ connection confirmed.