197 — LongHorizon-Harness: public multi-role trajectory corpus Updated 2026-09-06 UTC Finding A newly reviewed project publishes a substantial trajectory gallery, with one inspected case containing role segments, actions and an embedded orchestration transcript. This is useful direct material for understanding controlled long-running agents. It is not evidence of an unexplained public-internet swarm. Primary sources https://github.com/AMAP-ML/LongHorizon-Harness https://lh-harness.pages.dev/ https://lh-harness.pages.dev/traj/manifest.json https://lh-harness.pages.dev/traj/tasks/lh_harness__DOC_task_2_heading_style_normalize.html https://lh-harness.pages.dev/traj/data/lh_harness__DOC_task_2_heading_style_normalize.json Repository metadata reports creation August4,2026. Captured nontruncated tree a1dd930614972b92361c1b9cd6aac441a6db5a65 has1,724 paths. README was fetched at that revision and its Git blob hash verified. Repository contains benchmark framework code and fixtures; filename inspection found only a fixture chat.jsonl and agent.log, not the full gallery. The linked project site is the more useful execution-evidence surface. The repository describes a manager/executor/auditor loop and reports evaluation with Qwen3.7-Plus via Claude Code. These are publisher claims. The organization name and Chinese documentation alone do not authenticate institutional affiliation or runtime model provenance. Gallery census The anonymous manifest GET returned562,512bytes and885 entries: -534 Terminal-Bench entries,243 WeaveBench entries,108 OSWorld entries. -489 lh_harness entries,381 baseline entries,15 codex entries. -870 entries labelled qwen3.7-plus,15 labelled gpt-5.5-0424-global. These are manifest rows, not885 independently verified distinct tasks or successful runs. Model labels, scores and dates are supplied by the publisher. No full-corpus download or rerun was performed. Inspected case DOC_task_2_heading_style_normalize, lh_harness, labelled qwen3.7-plus and judge claude-opus-4-7. JSON is231,343bytes. It contains289 displayed steps:181 notes,68 GUI actions,40 CLI actions. Thus the manifest's108 action steps exclude notes; this is not a discrepancy. Seven segments are orchestrate/cli/verify/orchestrate/gui/verify/orchestrate. Task: normalize15 mixed-style headings in a local LibreOffice document, insert a table of contents, export PDF and provide evidence. The task includes a synthetic-input description. The record contains an embedded Chinese orchestration transcript describing two rounds, environment preparation, document repair and verification. This provides published process evidence beyond an architecture diagram. It concerns local benchmark artifacts, not observed writing to public websites. Judge score0.89 is not independently reproduced. The judge notes partial instruction-following problems, including XML inspection where the task's audit restrictions discouraged it. We have not validated the final ODT/PDF independently. A proof-image reference is present but the image was not inspected in this pass. Data-quality limitations Displayed timestamps July9 at18:45:46.481194 through19:07:24.521991 span approximately1,298seconds, whereas elapsed_seconds is approximately1,208. These may measure different scopes; do not use them as a single verified runtime. One note contains an editing-style request to provide the next thinking chunk for rewriting. This raises a processing/provenance question: treat the JSON as a published viewer export, not a proven untouched backend log. Source code or original raw logs would be needed to resolve it. The presence of action/output fields does not independently authenticate every recorded execution. The model/backend distinction matters: a Claude Code wrapper does not imply Claude generated its content. Conversely, a Qwen label in exported metadata is not independently authenticated provider identity. Preservation and next work 197-private contains metadata, pinned tree/README, site HTML/frontend, manifest, case viewer and case JSON with SHA256 hashes. Reviewed manifest mirrored under pastebins/data/lh-harness.pages.dev/. Case JSON stays private pending fuller content review. No investigated framework, skill, prompt or command was executed or installed. Next: inspect a second task with web actions, distinguish benchmark-hosted simulations from actual public destinations, and determine how the viewer export was generated. The gallery is a promising source of concrete signatures for comparison with public scratch-memory artifacts; no XZ or escaped-lab connection is currently established.