OKF-DWP

Independent experimental publication. Not an official DWP service, benefits advice or an entitlement calculator.

View this version’s Markdown source

On this page

Monday delivery work log

Deadline: 21 September 2026, 10:00 Europe/London (09:00 UTC).

This is the retained Monday work log. For the later source-led DMG/ADM build, use the current results and gaps and implementation log. The dated service observations below remain historical; service publication status selects the latest recorded public receipt.

Starting point, 20 September at 22:18 BST

DWP main 8d3349e7be91112fdc81ada1c9dda5ccdfcef60d and Explorer main fc71d65b8f5cfc860d52afe98a5e45b88231f3e5 have passed their recorded canonical checks. The public service is 0.4.0. Earlier releases, trials and receipts remain immutable. Private .email.md remains outside Git.

Initial concurrent work and ownership

Owner Bounded work State
Semantic agent Household qualifications, missing paths and fixed-evidence trials Authored increment reviewed; six inputs frozen; five actual attempts retained, zero accepted answers; separate CLI compatibility investigation
Legal agent Statutory producer and public service update 20 units and 43 links implemented; service 0.5 candidate preserves four source versions
Explorer agent Ambiguity, bounded parallel reads and browser verification PR 128; exact-build acceptance refreshed; household journeys pass three local browsers
Integrator Separate learning website, contracts, builds, evaluation, independent review and protected publication In progress

The separate Explorer security-review checkout is untouched. All producer and release writes are serialised by the integrator; agents edit disjoint authored inputs. The task heartbeat resumes hourly through the deadline and checks active workers before starting work.

Findings requiring a correction to the initial diagnosis

All eight missed candidate pages are present in the existing index, but none was reachable from the resolved seeds. The node-budget flags on staff-005 and staff-025 do not establish why those candidates were missed: raising limits alone does not add a missing path. Six require evidenced authored relationships; two belong to separate unresolved meanings of SDA. Those meanings must remain ambiguous even when both branches can be inspected.

Release conditions

Re-run affected source, semantic, retrieval, budget and provenance controls. Retain new receipts and model trials separately; never overwrite earlier runs. Update the backlog and beginner documents with actual measured outcomes. Merge reviewed increments only after required checks, then verify canonical CI and any affected public journeys. Specialist acceptance, CPAG permission and physical Voice/PA rehearsal remain separate from work that agents can complete.

Next checkpoint

Integrate the three bounded increments, compare all 40 development questions, and rerun affected fixed-evidence claims against the new version. Report source closure separately from unresolved scope, legal applicability and specialist review. Freeze new scope at the deadline and record tested results and gaps.

Integrated candidate checkpoint

The old staff trial and its five paired cases remain untouched. The next trial uses a new immutable context/prompt freeze; both evidence and instructions change, so it is not a controlled attribution of improvements to one component.

Verification checkpoint, 20 September at 23:20 BST

Next: finish protected merges, publish and verify the four-version service, then retain any separately reviewed model experiment and update the handover. The user has explicitly asked that side questions do not interrupt this work.

Protected source merge, 20 September at 23:23 BST

DWP PR 15 merged as 5ec1a158b107e3932ffcd1cd3a75e31494ee07db after both required validation runs passed. The source candidate and model inputs retain their original 3ef0e786e9a18e76fa17c7d925ff509d6d6c9f84 identity. A separate follow-up branch carries new observations and documentation; it does not change the frozen source or trial. Canonical validation and the later service deployment remain separate release checks.

The learning website's actual HTTPS verification also passed: 109 responses (manifest plus 108 files) matched the exact independently generated build. The receipt retains 1,473,921 transferred bytes and no retries. This check reads documentation only and does not fetch corpus bodies or private correspondence.

Explorer merge and service review, 20 September at 23:30 BST

Explorer PR 128 merged as 9b5bfc5a273905d125581e4c9a3a56023c7632b1 after all required checks passed, including the full Chrome/Firefox/WebKit gate. Its Pages run is 35541865447; publication remains under verification. DWP PR 16 preserves the new local-browser, learning-publication and original five-attempt model observations, with independent integrity review.

Explorer PR 129 carries service 0.5.0 at 45fd5e02ede8fe7e306a69285c30e88f85cc3cfd. Independent review and 47 local tests pass. Four source versions work through both supported MCP client versions; 108 immutable source files are hash-checked. A corrected Source family label has a fresh integration receipt, and the earlier candidate observation remains retained. Required CI, deployment and actual public SDK/browser checks are next.

The original rejected model outcomes remain unchanged. Separate small diagnostic calls identify current wrapper formats, while an independently checked lossless projection preserves every original context field and reduces the substantive input packages by 12–15%. Neither finding establishes better answers or speed.

Independent review of the two new receipt checkers found unbounded or symlinked local reads. The correction adds pre-read size/type checks and parent-directory checks, with ten passing integrity and tamper controls. Recorded browser and website observations remain unchanged; both inventories still verify.

Model successor preparation

Independent review and the integrator's replay pass 40 Monday Python controls and all six exact dictionary round trips. The new protocol keeps the same 240-second execution limit and default subscription models, with no paid API, tools or configuration changes. It first tests both unknown-term controls; substantive calls depend on those observed results. The original five rejected or timed-out attempts remain unchanged.

The one-off reasoning diagnostic initially retained an opaque tool identifier as a dynamic field name. Its public artefact is now an explicitly labelled privacy projection with the original digest; the original remains outside this public repository. The successor uses strict event shapes and does not reuse the diagnostic's generic redactor as an acceptance rule.

Live service checkpoint, 20 September at 23:50 BST

DWP source PR 15 passed canonical validation 35541533112 and learning-site publication 35542299240. Explorer engine Pages run 35541865447 passed. Service PR 129 merged as d538de99e6567633204253cd88b87cbe325ac39a and its canonical CI 35542085493 passed. Sites version 9 then deployed that runtime successfully at 22:36:48 UTC without changing the public audience.

The actual remote SDK run passed seven full-package cases and four compact reconstructions across four source versions: 93 successful requests, no retries. Three public browser evidence journeys then reconstructed the same care-home package exactly: 35 records, 50 relationships, insufficient evidence, no AI answer. Chrome and WebKit passed strict console checks; Firefox retained two hosting-cookie warnings. The 90-request browser run made no retries. The new release explicitly records historical browser journeys as not run.

The service-only Explorer merge exposed a publication gap: its nested guides passed the selected PR jobs but Pages rejected missing top-level documentation and changelog updates (35542056363). A separate correction updates both and moves lockstep checking ahead of CI impact selection. It changes no runtime bytes and retains the failed run. Publication of that correction remains open.

Both successor model controls passed event recognition and mechanical abstention. The separately frozen protocol is bound to source commit 78a8beea97242d646eb9159860dea190ca5e2998, manifest faefc7f42282c2f8dfdee119ff6d30f26b537b73b43b1b9ed133c3befbef4678. Only the first substantive pair is authorised at this checkpoint. The original five unsuccessful attempts retain their original outcomes.

Evidence and model checkpoint, 21 September shortly after midnight BST

The public full Reader passed a separate real Chrome journey against 270 unique immutable source files and all 21 application files. Concept filtering, the statutory graph, requested-version versus audit dates, the care-home heading and both unresolved SDA meanings were observed. The 16.6-second cumulative journey had no console or network errors. Its first harness reset-state assumption failed and remains retained. Both context packages remain insufficient and truncated. The public/local comparison preserves the different binding URL and derived context identities rather than asserting equal package IDs.

The successor model experiment is stable: nine attempts, seven parser-accepted responses, six mechanical passes including both empty-evidence controls. It retains 14 substantive claims and 21 citations, one defective source locator, a Claude formatter rejection and a Claude timeout. Three remaining Claude cases are explicitly held; no successful substantive pair is claimed. Model critique records scope and citation concerns separately from the mechanical results.

A further source review found that the care-home partner qualifications already exist in the captured corpus and authored semantic index. The small package removes them under byte pressure while retaining an authored summary. The task profile does not yet declare those qualifications required. The next bounded work separates a source-backed dependency correction from a generic allocation change: no DWP-specific routing in Explorer, no new acquisition claimed, and no rewriting earlier trials. Existing no-pressure context results must stay exact.

The presenter now has a Voice rehearsal sheet and a verified-evidence fallback. Actual Voice tool access and room audio remain untested. No account, microphone permission or audio setting was changed.

Editorial check: the eight affected overview, handover and verification documents were checked with Explorer's British-English checker. Its six flags all concerned judgment or judgments in references to judicial decisions. These intentionally retain the UK legal spelling; none is an American-English prose substitution.

Publication and review checkpoint, 21 September at 00:20 BST

DWP checkpoint c19446c1 is pushed to PR 16. It includes the current service and public Reader observations, frozen successor trials, Voice rehearsal and updated beginner documentation. Private correspondence and unrelated research remain outside the commit. Explorer PR 130 passed all checks, including the full browser suite, and merged normally; canonical CI and Pages are being checked separately.

A separate offline critique checker verifies seven retained answers, 14 claims, 21 citation diagnostics and four additional selected records. It preserves the invalid source-locator pair and rejects evidence or authority changes. Eleven tamper controls include a check that changing a model opinion does not make it machine-validated semantic truth. No frozen experiment file was changed.

The next source and engine improvements have separate working trees. Qualification dependencies will be declared in DWP authoring; Explorer will allocate bounded context using those general dependencies. Recomputing missing declared dependencies after trimming deliberately corrects an existing diagnostic gap: an omitted dependency must remain visible even if its relationship was trimmed. Historical observations and model trials remain immutable.

Initial qualification source checkpoint, 21 September

The isolated source candidate adds eight required-support assertions and two explicit qualification profiles. It retains 901 records, now with 1,435 semantic assertions. All seven declared support pages already existed; no source text was acquired or edited. Producer tests and Explorer input validation pass. The combined Reader build preserves 20,044 records and now has 21,195 relationships. Its existing requirement labels apply in both directions; no new vocabulary or display fallback is necessary. Context allocation and publication are separate remaining checks, and no one has closed the 203 outstanding obligations.

Component qualification follow-up, 21 September

The bounded component review led to two narrower definitions and seven additional support relationships. The housing summary now describes the former-home treatment and Housing Benefit expenditure that may be payable. The temporary-residence summary keeps its no-partner opening. Five housing-cost pages and two temporary-residence pages are required support for their respective summaries. Staff 012 and 013 add the housing-cost closure; the permanent-care-home questions do not unconditionally require the temporary-residence branch. The broader no-partner and severe-disability overview support options remain explicit follow-up work.

The current semantic candidate has 901 records, 1,442 assertions, 15 support dependencies and two qualification profiles. Each of Staff 012 and 013 declares 20 required identifiers and 17 paths; its five open obligation identifiers are still absent evidence. All 203 obligations and original candidate identifiers remain unchanged. Its snapshot is dwp-staff-semantics-fc2ad1545adcbeeb2243 and index SHA-256 is 7ffc9d00e71fef6aed5531373510df82998123e89fdb28091cfb82384adf2876.

The regenerated combined Reader retains 20,044 records, 513 source PDFs and 19,090 source pages, with 21,202 relationships. Its snapshot is dwp-combined-9b85571294b473b9fe6e. The 20 selected statutory units remain separate from the manual counts. Current producer and Reader outputs are reviewable engineering artefacts; the earlier eight-edge checkpoint above retains its original counts. Joint context allocation, exact-version publication and answer quality remain separate checks. No frozen source, model-trial package or browser observation was rewritten.

Verification for this follow-up: all 31 staff producer controls, 20 combined Reader controls and 15 backlog controls pass. The Reader checks cover all 15 requirement edges, their four source-concept groups, both navigation labels, provenance and unreviewed authority. The backlog projection matches its register. Both producer --check commands pass; the combined check reproduces all 4,758 declared output files exactly. These offline checks do not establish a later browser or AI-answer result.

Joint qualification verification and ledger update, 21 September

Joint receipt SHA-256: 45fb3ce7b768833eb6f94b6461f1531f56293de419665bd5adaf606d9bfdc3ac. The evidence-path improvements do not retrospectively repair or regrade any model answer. No new model or public HTTP calls were part of this comparison.

Continuing independent increments

Two follow-on implementations are running in isolated checkouts. The disability addition work narrows the summaries, adds the two captured treated-receipt pages and reviews their conditional dependencies. It is tracked as DWP-BL-007.component-qualification-expansion; the current qualification release and its fixed comparison are unchanged.

The service work adds explicit assembler identity and bounded historical replay under DWP-BL-008.engine-replay. The existing progressive evidence-read milestone retains its recorded completion; the aggregate item is now in progress because this additional work has started. Public retained evidence resources have their own pending package. Neither workstream has deployed a new service or rerun the frozen model trials.

Qualification publication and disability follow-up, 21 September at 01:35 BST

Explorer PR 131 merged as 9cf67adb12cd39c6c008336de92929a08972a440. Canonical run 35546895467 and Pages run 35546875345 both passed. A fresh public Chrome observation at 01:23 BST bound the deployed application manifest 9fc8cb1bbf10e4e5182efd69d56f2b5ed39a2e6ecf942dce64357c4a529d1ce8 to the qualification source 7f9feb9634e3d94004853b838462aca132c505a5. Six journeys passed, with 270 exact corpus files, 294 responses and no console/network errors. The care-home package has 62 records and 124 relationships; SDA retains both unresolved meanings. This is a new observation, not a relabelled earlier receipt.

DWP PR 17 merged as 91b9907836c8a340d974dd958969f4f8cbb3a0c6 after both required validation runs passed. A merge commit preserves the source/evaluation commits. Canonical validation 35547851197 and the subsequent learning-site publication are separate checks still under observation at this checkpoint.

The legal agent authored and the semantic agent independently reviewed the disability increment. Root retained both receipt qualifiers identified in review and the stricter existing housing wording control when integrating the two overlapping edits. All 46 semantic and 20 Reader controls pass. The semantic index has 903 records, 1,464 assertions, 98 selected pages and 29 support dependencies. All 203 obligation identifiers and statuses remain open; two source-closure labels are intentionally clearer.

Root rebuilt all 4,758 combined outputs and the current 40-case evaluation, then verified exact regeneration/replay. The combined Reader has 20,044 records and 21,224 relationships. The separate disability comparison uses source 8ea4465a4cb5a867d82c87e635f2ef1d1df18d8a, two archived engines, two budgets and all 40 question occurrences. Its 320 assemblies and exact replay pass, with 14 admission/denominator controls. The current engine retains 497/497 paths at 512 KiB but 407/497 at 256 KiB. Candidate overlap remains 177/177 at both sizes; it cannot establish qualification coverage. All contexts remain insufficient. The 64 KiB care-home request explicitly refuses without evidence.

The next partner-component source work is isolated and tracked separately. The service's explicit-engine replay implementation has passed independent runtime review after a shared-deadline correction; its final build/browser checks and deployment remain separate. A new direct-JSON subscription trial protocol has been proposed but no calls or new answer freeze have occurred. All earlier model failures, trial receipts and the live 0.5.0 service are preserved.

Publication checkpoint, 21 September at 01:50 BST

The qualification merge 91b9907836c8a340d974dd958969f4f8cbb3a0c6 now has successful canonical validation 35547851197 and learning Pages publication 35548811691. The retained learning-site receipt still identifies the earlier 7815b17b… release; a successful hosting workflow is not a new HTTP observation.

Root reviewed and ran the new qualification public-receipt checker and all 14 controls. It verified 273 immutable source inputs, 21 application materials and 90 whole-evidence occurrences across the care-home and SDA packages. Every selected and returned required path is retained. Separately, 42 and 35 returned relationship rows point to 37 and 31 unselected targets with explicit budget omissions. The documentation now explains that those references do not supply the omitted evidence. All 11 original browser files remain unchanged.

The partner-component increment passed a separate source review: 38 staff and 14 household controls pass; the 903-record index has 1,482 assertions and 39 support dependencies. All 520 pre-existing evidence records and the identities, categories and statuses of all 203 obligations are preserved. Four conditional labels are intentionally clearer. Combined projection, bounded context retention and public delivery for that next increment remain separate work.

CI ordering correction

PR 18 run 35548995209 failed because the new receipt test executed before CI fetched its immutable 7f9feb96… source. The local checkout already contained that commit, so its 396 passing tests did not expose the shallow-checkout dependency. The workflow now fetches that exact source before the unit suite; the later comparison step reuses it. No test, receipt or source is weakened or rewritten. The failed run remains visible; the corrected commit needs a fresh required CI pass before merge.

The subsequent run 35549578095 reached all 396 tests and exposed a second portability problem: six temporary-directory cases used macOS /private/tmp, which does not exist on Ubuntu. The test now resolves Python's platform temporary directory before constructing its isolated fixtures. The symlink checks and frozen checker/receipt bytes are unchanged; all 14 focused controls and the offline receipt check pass locally. A new Linux CI run remains required.

Partner retention checkpoint, 21 September shortly after 02:00 BST

The reviewed partner source is frozen at 7e5fdb9b906052914b307c17c0fd19feb2d008a7. Root rebuilt and checked all 4,758 combined outputs: 20,044 records, 21,242 relationships and snapshot dwp-combined-a769a75d999ffd48afb0. All 402 Python controls, the 40-case current staff evaluation and exact replay pass. All 203 obligations remain open.

The separate partner comparison runs 320 assemblies and exact replay. The legal agent independently checked 380 immutable input bindings, archived code, the 497-to-585 path denominator and eight exact focus-package reconstructions. Fourteen admission/census controls pass. At 512 KiB the current assembler retains all 585 declared path occurrences; at 256 KiB it retains 449. Candidate overlap is 177/177 at both budgets. Staff 012/013 retain 93/93 paths at the larger size and 25/93 at the smaller size. Missing qualifications remain visible.

The learning site at revision 91b99078… passed a new actual HTTP observation at 01:53 BST: 123 matching responses, 1,777,711 bytes, 120 HTML pages and 2,382 internal links. The executed verifier and receipt are retained separately while the next publication record is prepared. The disability candidate df352daa… also passed six actual public Chrome journeys against application manifest 9fc8cb1b…, with 270 distinct corpus files and no console/network errors. Its own source/receipt identity is separate from the earlier qualification check.

The service replay verifier passed independent review after stricter error-only response checks and pre-open special-file checks. All 77 service controls pass; no new public service call has occurred. The direct-JSON model runner has 24 offline controls and is under independent review; no provider has been called. Ignored-person discovery has produced a concrete source proposal, with its normal-residence distinction and remaining statutory/judgment gaps explicit.

Ignored-person coverage and capacity checkpoint, 21 September at 08:10 BST

The separately reviewed ignored-person source c44bc3a1… adds two conditional concepts and eight already captured source pages. Root rebuilt and checked all 4,758 combined files: 20,046 records, 21,286 relationships and 19,912 resources. All 411 Python controls and the 40-case current evaluation/replay pass. The original 520 evidence records and 203 obligation objects remain unchanged.

The separate frozen comparison preserves 320 assemblies and exact replay. Eighteen admission/census controls pass. At 512 KiB the new source retains 177/177 candidate occurrences but only 647/765 declared path occurrences; at 256 KiB these fall to 171/177 and 393/765. The care-home cases lose one of seven household support pages even at 512 KiB. The byte census reproduces each measured package hash: Staff 012 contains only 45,608 JSON bytes of source evidence text, while requirements and relationships occupy 124,383 and 170,815 bytes. This identifies representational overhead for investigation; it does not justify silently dropping dependencies or declaring an answer complete.

The source modelling increment is separately recorded complete. Its capacity regression has a named open work package. The Monday service and model trial remain fixed to the partner source 723bcc5b…. Work was interrupted overnight by usage limits; on resumption, source PR 19 passed and merged normally, and the reviewed service passed all protected checks. New live observations and model attempts are separate from these offline measurements.

Retained later public observations, 21 September

Root preserved the 01:53 BST learning-site check for exact source 91b99078…: 123 HTTP 200 responses, 120 HTML pages and 2,382 internal links. The separate 01:55 BST public Chrome observation of disability source df352daa… passed six journeys and retained the complete packages. Its offline admission checks 273 immutable source inputs, 71/71 returned required care-home paths and 12/12 returned required SDA paths; frontier relationships remain separately reported. Six new local checker-admission controls cover altered bytes, file bounds, symlinks and non-regular files. These observations do not attest the later partner source or the undeployed MCP successor. Historical files and .email.md remain untouched.

Versioned public delivery and retained examples, 21 September at 08:30 BST

Root deployed service 0.6.0 once, retaining the exact Worker, source and hosting identities. The first public SDK check delivered all 11 evidence cases but failed its comparison of different SDK envelopes. A reviewed verifier-only correction preserves complete tool rows and exact schema/trust checks. The fresh run passed 121 requests, including all nine compatible source/engine pairs, an empty control and exact historical replay. The failed run remains unchanged.

The reusable offline archive exporter passed independent review, eleven controls and a local seven-check Chrome journey. The DWP publisher passed 23 controls; its separate public byte verifier passed thirteen. Root exported three approved actual public packages into 225 files/2,766,291 bytes and admitted their exact Git-bound inputs. Public Pages and browser verification remain separate gates.

Both direct-v3 subscription controls completed within bounds but their parsers rejected undocumented-for-that-run metadata. The receipts preserve the failure categories and unknown tool census. No substantive call or retry followed. Installed-client schema inspection is informing a separately reviewed successor; the frozen v3 inputs remain unchanged. This is not a successful paired answer-quality or affordability result.

The first integrated 457-test run exposed one draft-only test that still required the protocol to be pending after the approved freeze. Its lifecycle assertion now requires the committed freeze while separately exercising refusal of a synthetic pending protocol. Frozen executable inputs and actual trial outputs are unchanged.

CI temporary-directory portability correction, 21 September

PR 21's eight retained direct-v3 observation controls failed on Linux because the test fixture required the macOS-specific /private/tmp directory. The test now uses Python's system temporary-directory default and resolves the resulting path before applying the existing strict symlink checks. All eight controls pass locally with both the ordinary environment and an explicit /tmp symlink on macOS. Linux CI must confirm the candidate after publication; these local checks do not claim a remote CI pass.

A scoped scan of test files changed in the previous 15 commits found no other hard-coded platform temporary-directory requirement. The other guarded Python fixtures already resolve tempfile.gettempdir(), while the browser controls use Node's tmpdir(). Literal private paths in older negative tests are synthetic rejection/redaction inputs, not directories to create. Only the fixture and these documentation notes changed; frozen executable inputs and recorded provider outcomes remain untouched.

The next push-run failure occurred during fixture cleanup: a background Git pack directory changed while Python removed a synthetic repository. Fixture repositories now disable automatic garbage collection and maintenance locally; real repository configuration and publication validation are unchanged. The failed run remains visible in CI.

Paired direct responses, 21 September

The separately reviewed v4 freeze preserved the complete v3 packages, prompt and answer schema. Its parser adds documented installed-client metadata recognition, bounded structural diagnostics and strict numeric usage fields. Thirty-nine controls and independent review passed before any provider call.

Both actual empty controls passed and abstained. The root then made exactly one Staff 012 call per subscription client: both passed mechanical checks with complete event census and zero observed tools. The two answers contain six claims and seven exact citations. Both distinguish the whole award from additional amounts and decline a whole-award conclusion. Independent agent claim review is separate; specialist acceptance, model identity parity, comparative accuracy and affordability remain unestablished. No trial was retried or earlier failure overwritten.

The independent agent review binds four attempts and 21 selected qualification records. It finds the six main scoped statements traceable, but records omitted treated-receipt/transitional exceptions, an overstated gap and further citation needs in one response. The answers remain unchanged and BL010 stays open for qualification repair and human assessment. The integrated local suite passes all 507 Python controls; public and protected-main checks remain separate gates.

Service 0.6.1 and native client checks, 21 September 2026

The question-schema patch was published as Sites version 11 at 11:52:32 BST. The reviewed runtime 1420c316… and merged Explorer commit 68743984… share complete Git tree 67eb52f7…; the original build identity is preserved. An initial archive save was rejected before a version was created because of the entrypoint layout. The corrected final archive used unchanged Worker bytes. The 0.6.1 release record links the exact archive, storage, runtime and publication identities.

The new actual SDK run passed 11 cases in 121 requests, receiving 10,322,722 bytes from 11:53:33 to 11:55:35 BST. It includes all nine allowed source/engine pairs, the historical care-home package and an empty control, without retries, model calls or full-package tool calls. Every complete package hash matches the preserved 0.6.0 run. The larger Staff 012 package remains insufficient.

After refresh, ChatGPT settings advertised the corrected pattern, three tools, five sources and two engines; permissions were unchanged. The same previously rejected native unknown-term control then succeeded. The exact Staff 012 native smoke check also returned, but its 16 KiB budget retained no records and reported an explicit byte-budget omission. It does not establish useful answerability. The existing Codex task still exposes only full ask_okf; compact-client and nested Data Agent access remain unproved. BL008.client-connection stays in progress, and no new AI-answer or Voice acceptance is claimed.

The status checker was also exercised against real drift: new receipts with the old 0.6.0 selection failed with “A newer successful deployment is not represented”. Pinning the receipts at immutable commit a13a291f… produced the 0.6.1 status and passed check mode plus 19 controls. The before/after client records and all older SDK failures remain separately retained.

Abroad semantic audit and reusable question diagnostics, 21 September

Root traced the reported five-page result to the 19 September pre-fix observation, then ran ten offline packages against exact source and engine identities. The published 723 source resolves the exact abroad question at 32 KiB, retaining two records and one relationship; overseas paraphrases remain unresolved. This is a local replay, not a new public or model observation.

Separate agents own the additive DWP international graph and its source/path tests, the reusable Explorer question-scaffolding classifier, and independent whole-passage review. Root owns integration, backlog, documentation and release boundaries. The existing four international IDs are reused; no frozen source, engine, receipt, trial or private email is rewritten. The audit names cross-benefit and graph-budget work still outstanding. The Data Analytics guide proposes a controlled semantic-proposal exercise; it is not a completed new model trial.

The bounded source repair passed independent agent review, 67 semantic tests, 21 combined Reader tests and the current 40-case replay. Its separate comparison retains 24 complete packages and 10 negative controls. Four Pension Credit cases retain 7/7 whole pages and 14/14 required paths at 512 KiB; the general 32 KiB result regresses to one concept and no source evidence. The regression is explicit in the retained report. Public-service admission and generic capacity repair remain separate; all 203 obligations stay open. Explorer's shared-classifier change passes 607 tests and has its own reviewed PR; frozen MCP engines are not rewritten.

At the owner's request, a separate ChatGPT Work task returned a frozen-file semantic review and proposal summary for source 723bcc5b…. Its reported workflow used local Git and file/PDF inspection after raw web reads failed; it reported no callable Ask OKF tools or separate nested Data Agent call. The local task inspected the returned messages and proposal summary, not the two complete cloud artefacts. Their reported hashes remain unverified locally. Novel proposals await full artefact import and independent source review; the method guide records this boundary rather than claiming native MCP acceptance or an accuracy benchmark.

Case-level wider semantic audit, 21 September 2026

A read-only audit inspected all 40 retained current after packages (39 distinct questions) and verified their decoded hashes against the evaluation at source commit c203a4bd621e57c99273b3933df0207e101c5a85. The packages use the pinned c4f2de0a… engine, a 512 KiB limit and semantic index SHA-256 92a8871b8f1f2e51f1feace0b1f57c67dfd0ddcb0434cf4574795e5942e95fd6. No assembly rerun, model call or public-service test was made for this audit.

The wider-work section now names the remaining work: 27/38 ADM-containing packages have no ADM relationship path (27/40 overall); 24/40 activate multiple profiles; and 23/40 retain unresolved tokens, including both domain constraints and ordinary wording. It distinguishes duplicated and broad profile triggers from missing evidence. Only cases 012/013 lose declared required paths: 118/776 path occurrences across the 40 packages, or 62/463 after deduplicating identical paths within each package. Separately, 15/40 have 86 support-dependency diagnostics, all pointing to records present in the index. Existing backlog packages cover these findings. All 40 packages remain insufficient and all 203 obligations stay open. These checks describe recorded selection and diagnostics, not specialist acceptance or a new legal answer-quality result.

The local browser review of the committed learning website exposed two stale learning-path descriptions: the returned Data review still read as unrun, and the dated 0.6.0 baseline was called current. Both are corrected. Mutable service status now points to the generated, receipt-checked record; historical counts retain their own dated observation. No service deployment was performed.