Multi-agent work log: 19 September 2026
User-facing changelog · Backlog · Retrospective
Branch: codex/evidence-review-navigation. State: dated observations;
integration and validation status are linked below.
This ledger records responsibilities, decisions and handovers. The changelog
summarises changes useful to readers; it is not a dump of agent messages or Git
commits. A completed subtask is not a merged or deployed release.
Work ownership and dependencies
| Task ID | Owner role / workstream | Owned files or interface | Dependencies | Current handover |
|---|---|---|---|---|
| DWP-WL-001 | Integration / root | Reusable evidence-delivery contract, remote service and human review route | Existing immutable context identity; DWP-WL-002/003 interfaces | Public service 0.3.1 and exact SDK acceptance recorded; v7 historical and full-corpus functional checks passed, strict console gates failed; v6 preserved (DWP-BL-023) |
| DWP-WL-002 | Corpus semantics review | domain-profile/navigation/, navigation producer and staff-review packs; narrow additive contracts |
Frozen DMG/ADM bytes and existing concept IRIs | Producer checks and exact local/public DMG Reader observations recorded; ADM Reader remains BL024 |
| DWP-WL-003 | Explorer navigation | Generic conceptual navigation and full-text human review consumer | DWP-WL-002 manifest; reusable consumer contracts | Exact local and published Chrome navigation journeys passed; service hosting-console failure remains separate |
| DWP-WL-004 | Methodology and evaluation | Methodology/template, ontology map, backlog, retrospective, answer-trial harness; README/CHANGELOG integration | Actual namespace audit from DWP-WL-002; fixed public contexts; existing Claude subscription | Methodology/backlog handed over; fixed trials and actual local compact client replayed; strict failures preserved; human review pending |
| DWP-WL-005 | Integration and independent review | Exact candidate review, CI, PR, release and public observations | All affected workstreams, current documentation and checks | Explorer PR 125 merged, Pages and public DMG navigation passed; DWP PR 10 records integration and exact check/merge status |
Public artefacts only. Private correspondence, existing research scratch files, credentials, account identifiers and local browser state are excluded.
Decisions and rationale
| Decision | Rationale and limit |
|---|---|
Preserve ask_okf; add ask_okf_manifest and read_okf_evidence |
Existing clients retain their full-package behaviour. Smaller delivery is additive, not a change to evidence authority. |
| Recreate approved context for each evidence read | Question, version and original budget must yield the requested context ID before any record is returned; no stored question session or arbitrary source URL fetch. |
| Exact text slices with offsets, digests and continuation | A consumer can reconstruct evidence without silent summaries. Offsets are UTF-16 units as defined by the shared implementation. |
| Human review link carries a replay recipe in the URL fragment | The HTTP request does not carry the fragment as a query. The page asks the user to recreate the evidence; it does not reproduce a stochastic AI answer. |
| Formal MCP resource reading deferred | Prove the read-only tool interface and compatible HTTPS review link first; do not claim untested client support. |
| New navigation is additive | Preserve frozen source and old semantic builds. Topic assignment is literal discovery, not appliesTo or a legal conclusion. |
| Claude trial is separate from engineering evaluation | Exact citation checks, legal entailment and human usefulness are different judgements. No model output becomes gold by generating it. |
Checks actually recorded so far
-
Claude Code help was inspected before use. An isolated, tools-disabled public probe failed authentication in the sandbox and succeeded at the authorised normal host boundary. No login, installation or model override was performed.
-
The fixed-package harness has six passing controls for source identity, invented citations, altered quotes/locators, false sufficiency and no-evidence abstention. These test the checker, not the correctness of any model answer.
-
Actual baseline trials use three retained public packages and save their exact input/prompt/output hashes. Custody produced 12 claims/31 exact citations; the no-result control abstained with no claims. Abroad produced 10 claims/14 citations, with 10 strict quotation failures diagnosed as whitespace-only changes and four exact quotations. The strict failures remain recorded. Results are recorded in the answer-review directory; independent human review remains pending.
-
Namespace use was checked against the retained semantic shards and pinned context; see ontology use. Declared ELI/OWL prefixes were not promoted to implemented ontology claims.
-
The reusable service's compact-delivery verifier has eight passing offline controls: exact reconstruction and seven corruption rejections. It preserves the earlier complete-package comparisons and adds schema, annotation, catalogue, provenance, slice-hash, byte-bound and fail-closed checks. A live receipt is now recorded for the deployed 0.3.0 service.
-
The public backlog checker has eight passing fixture controls, including dependency cycles, private/escaping paths and a misleading prose status. CI now replays the three retained answer trials and their quotation diagnosis offline, without model calls or silently turning strict failures into passes.
-
An actual Claude local MCP retry made seven verified compact-tool calls and read five complete values, after a zero-call safe-mode attempt generated an invented catalogue. Both observations are retained in the compact-client report. Five offline controls preserve that failure and reject corrupted source/status/size values. No public deployment or model-answer acceptance is inferred.
-
The navigation browser observation records the separate exact local candidate's conceptual facets and journeys; it does not replace the dated public deployment receipts.
Later owners should append concrete test commands/results and receipt links, not change “in progress” to “passed” from implementation alone. Preserve failed attempts. Before handover, state modified files, exact inputs/consumer revision, tests run, remaining checks and any active process. Before publication, reconcile this ledger, the backlog and the relevant changelog entry with the actual candidate.
Public service handover
Hosting and
SDK verification bind runtime
169b8c387a29435d39dc31cbb2066376d84b39a6. The SDK completed at 20:21:55 UTC
on 19 September. It retained all three original full-package comparisons and
verified the smaller catalogue/read path, full reconstruction and three rejected
invalid reads. The five-minute guide keeps this
public-service result separate from local navigation, local Claude, the earlier
ChatGPT call and any later public-browser observation.
Public browser gate: hosting conflict retained
The public run summary
records the historical custody profile across Chrome, Firefox and WebKit. All
functional assertions in the nine cases passed before their final console
check. All nine overall tests failed because the strict no-console gate
reported the host's injected Cloudflare challenge loader, blocked by the unchanged
CSP; Firefox also reported invalid-domain __cf_bm cookies. The earlier local
9/9 result remains separate. That suite does not establish full-corpus browser
behaviour or AI-answer quality.
A separate public full-corpus Chrome observation
at 20:38 UTC made four real tool calls: catalogue, exact source text, provenance
and diagnostics. It displayed six records and the insufficient 2cdfa5fe…
abroad context; rendered text and metadata for ADM C2 page 18 and complete
diagnostic hashes matched the SDK. Functional checks passed, while the strict
console gate retained the host CSP error. This is hosting version 6, runtime
169b8c387a29435d39dc31cbb2066376d84b39a6, before the 0.3.1 replay-link
correction. No other browser engine or AI answer is covered by this observation.
DWP-BL-023 assigns a bounded hosting/CSP integration follow-up. Its acceptance requires supported hosting changes, preserved CSP and controls, and fresh exact functional and console observations. No inline-script allowance, console-error suppression or hosting-control bypass was introduced. SDK integrity and the local Claude evidence observation remain separately recorded.
Final review correction and service 0.3.1
Independent review found that changing or resubmitting the question could leave
an earlier replay link visible. The correction clears that link on question or
source changes and on resubmission; twelve local journeys across Chrome, Firefox
and WebKit passed. The 0.3.1 deployment
at 20:47:55 UTC and SDK receipt
completed at 20:48:50 UTC bind runtime
8493b323ca664e645a2548ebb48bf7917d7f6eb1 and hosting version 7. The SDK again
verified all three unchanged full packages and compact reconstruction. Earlier
v6 browser failures remain preserved.
The v7 historical-profile suite then exercised four journeys in each of Chrome, Firefox and WebKit, including question A → B and the replacement replay-link identity. All twelve functional sequences passed before the final strict console assertion; all twelve overall tests failed because the host CSP/cookie errors remain. A separate v7 full-corpus Chrome journey made four real calls and matched displayed source text, provenance and complete diagnostic hashes to the SDK's six-record insufficient abroad context. Its strict console check also failed. DWP-BL-023 remains open; no AI-answer or clean public-browser acceptance is claimed.
The final assembled-site check also exposed a documentation-cache defect: transitive linked service documents were rendered but absent from the cache fingerprint. The narrow build-script correction includes the existing bounded reading closure. Forty focused build/cache tests and the complete site build passed, including 340 reading pages and 7,449 internal references. Source evidence, service runtime and the application build were unchanged by that fix.
Integration checkpoint
Explorer PR 125 merged
as 8a38d4bfe07a6797deacc148a2859911d8e2a68e after the
required CI run
passed: 552 app tests, 493 Python tests, browser suites of 348 and 78 tests,
37 service tests and 12 local service-browser journeys. The
Pages deployment
then completed successfully at 21:15:34 UTC. The separate
public Chrome navigation observation
passed at 21:16:55 UTC against that merge and DWP corpus commit
0c59f476602a49d4293e46112a0ca1ebf1486ff0. It verified 21 application files and
188 observed corpus files over actual HTTPS, with no response substitution.
Benefit, circumstance, topic and authored-concept selections matched counts
74/343/268/1 across Reader, Graph and Timeline; capture/source roles and the
80-group Timeline limit were visible. No console errors or targeted axe
violations were found. Whole-application accessibility and ADM Reader coverage
are not established by this DMG observation.
DWP PR 10 records integration of the evidence and documentation; consult the PR for its exact validation and merge status. The observed model trials await specialist review. All 40 staff-question occurrences and all 43 full-corpus evaluation cases remain insufficient; this integration checkpoint does not establish complete evidence profiles or benefit decisions.