Monday delivery work log
Deadline: 21 September 2026, 10:00 Europe/London (09:00 UTC).
This is the retained Monday work log. For the later source-led DMG/ADM build, use the current results and gaps and implementation log. The dated service observations below remain historical; service publication status selects the latest recorded public receipt.
Starting point, 20 September at 22:18 BST
DWP main 8d3349e7be91112fdc81ada1c9dda5ccdfcef60d and Explorer main
fc71d65b8f5cfc860d52afe98a5e45b88231f3e5 have passed their recorded
canonical checks. The public service is 0.4.0. Earlier releases, trials and
receipts remain immutable. Private .email.md remains outside Git.
Initial concurrent work and ownership
| Owner | Bounded work | State |
|---|---|---|
| Semantic agent | Household qualifications, missing paths and fixed-evidence trials | Authored increment reviewed; six inputs frozen; five actual attempts retained, zero accepted answers; separate CLI compatibility investigation |
| Legal agent | Statutory producer and public service update | 20 units and 43 links implemented; service 0.5 candidate preserves four source versions |
| Explorer agent | Ambiguity, bounded parallel reads and browser verification | PR 128; exact-build acceptance refreshed; household journeys pass three local browsers |
| Integrator | Separate learning website, contracts, builds, evaluation, independent review and protected publication | In progress |
The separate Explorer security-review checkout is untouched. All producer and release writes are serialised by the integrator; agents edit disjoint authored inputs. The task heartbeat resumes hourly through the deadline and checks active workers before starting work.
Findings requiring a correction to the initial diagnosis
All eight missed candidate pages are present in the existing index, but none was reachable from the resolved seeds. The node-budget flags on staff-005 and staff-025 do not establish why those candidates were missed: raising limits alone does not add a missing path. Six require evidenced authored relationships; two belong to separate unresolved meanings of SDA. Those meanings must remain ambiguous even when both branches can be inspected.
Release conditions
Re-run affected source, semantic, retrieval, budget and provenance controls. Retain new receipts and model trials separately; never overwrite earlier runs. Update the backlog and beginner documents with actual measured outcomes. Merge reviewed increments only after required checks, then verify canonical CI and any affected public journeys. Specialist acceptance, CPAG permission and physical Voice/PA rehearsal remain separate from work that agents can complete.
Next checkpoint
Integrate the three bounded increments, compare all 40 development questions, and rerun affected fixed-evidence claims against the new version. Report source closure separately from unresolved scope, legal applicability and specialist review. Freeze new scope at the deadline and record tested results and gaps.
Integrated candidate checkpoint
- Learning-site PR 14 merged as
3f72ebc30128c2a3171951050a566d3ed8db7c16; canonical validation and Pages publication are tracked separately. - Household semantic index: 901 records, 1,427 assertions and 4,581,721 bytes; 51 authored concepts, 96 selected source pages and 20 selected statutory units.
- Forty-question replay: 176 of 177 known candidate occurrences; all 40 packages remain insufficient and all 203 named obligations remain open.
- Combined Reader: 20,044 records, 21,187 relationships, 19,912 resources, all 513 original PDFs and 19,090 pages. Statutory units are a separate denominator.
- Independent review corrected a missing verified metadata endpoint, an omitted statutory source-plane identity and the declared consumer version. The Reader keeps statutory dates separate from publication and links exact captured text.
- 309 Python tests pass; old browser observations still verify against 261 exact archived source files and explicitly do not attest the current candidate.
- New trial protocol and harness passed 23 independent offline controls after incomplete-stream and unrecognised-tool-event fixes. This was the checkpoint before the first new model calls.
The old staff trial and its five paired cases remain untouched. The next trial uses a new immutable context/prompt freeze; both evidence and instructions change, so it is not a controlled attribution of improvements to one component.
Verification checkpoint, 20 September at 23:20 BST
- The learning website is public at
https://chris-page-gov.github.io/okf-dwp/docs/learning-path.html.
Canonical validation
35540070101and Pages deployment35540739295passed. A browser check opened the learning path and glossary without console errors or page-level horizontal overflow. Exact HTTP publication checks are separate. - DWP candidate
3ef0e786e9a18e76fa17c7d925ff509d6d6c9f84is PR 15. Its branch validation passed; the separate pull-request run is still active. Explorer PR 128 at6f87c55d9152cd40e509141a61c9808e746e9013has passed the non-browser gates; its full browser gate is still active. - Three fresh local household journeys pass in Chrome, Firefox and WebKit. The receipts retain 41 artefacts, including two failed harness assumptions. The integrator independently ran the integrity verifier and seven admission/integrity controls. These are local candidate checks, not public deployment receipts.
- The six new model inputs are frozen against the exact DWP and engine commits. Five subscription attempts yielded zero accepted answers: two Codex event mismatches, two Claude timeouts and one Claude event mismatch. The frozen harness and all outcomes remain unchanged. Separate diagnostics will determine whether a new protocol is justified; they cannot retrospectively accept these rejected trials.
Next: finish protected merges, publish and verify the four-version service, then retain any separately reviewed model experiment and update the handover. The user has explicitly asked that side questions do not interrupt this work.
Protected source merge, 20 September at 23:23 BST
DWP PR 15 merged as 5ec1a158b107e3932ffcd1cd3a75e31494ee07db after both
required validation runs passed. The source candidate and model inputs retain
their original 3ef0e786e9a18e76fa17c7d925ff509d6d6c9f84 identity. A separate
follow-up branch carries new observations and documentation; it does not change
the frozen source or trial. Canonical validation and the later service deployment
remain separate release checks.
The learning website's actual HTTPS verification also passed: 109 responses (manifest plus 108 files) matched the exact independently generated build. The receipt retains 1,473,921 transferred bytes and no retries. This check reads documentation only and does not fetch corpus bodies or private correspondence.
Explorer merge and service review, 20 September at 23:30 BST
Explorer PR 128 merged as 9b5bfc5a273905d125581e4c9a3a56023c7632b1 after
all required checks passed, including the full Chrome/Firefox/WebKit gate.
Its Pages run is 35541865447; publication remains under verification.
DWP PR 16 preserves the new local-browser, learning-publication and original
five-attempt model observations, with independent integrity review.
Explorer PR 129 carries service 0.5.0 at
45fd5e02ede8fe7e306a69285c30e88f85cc3cfd. Independent review and 47 local
tests pass. Four source versions work through both supported MCP client versions;
108 immutable source files are hash-checked. A corrected Source family label
has a fresh integration receipt, and the earlier candidate observation remains
retained. Required CI, deployment and actual public SDK/browser checks are next.
The original rejected model outcomes remain unchanged. Separate small diagnostic calls identify current wrapper formats, while an independently checked lossless projection preserves every original context field and reduces the substantive input packages by 12–15%. Neither finding establishes better answers or speed.
Independent review of the two new receipt checkers found unbounded or symlinked local reads. The correction adds pre-read size/type checks and parent-directory checks, with ten passing integrity and tamper controls. Recorded browser and website observations remain unchanged; both inventories still verify.
Model successor preparation
Independent review and the integrator's replay pass 40 Monday Python controls and all six exact dictionary round trips. The new protocol keeps the same 240-second execution limit and default subscription models, with no paid API, tools or configuration changes. It first tests both unknown-term controls; substantive calls depend on those observed results. The original five rejected or timed-out attempts remain unchanged.
The one-off reasoning diagnostic initially retained an opaque tool identifier as a dynamic field name. Its public artefact is now an explicitly labelled privacy projection with the original digest; the original remains outside this public repository. The successor uses strict event shapes and does not reuse the diagnostic's generic redactor as an acceptance rule.
Live service checkpoint, 20 September at 23:50 BST
DWP source PR 15 passed canonical validation 35541533112 and learning-site
publication 35542299240. Explorer engine Pages run 35541865447 passed.
Service PR 129 merged as d538de99e6567633204253cd88b87cbe325ac39a and its
canonical CI 35542085493 passed. Sites version 9 then deployed that runtime
successfully at 22:36:48 UTC without changing the public audience.
The actual remote SDK run passed seven full-package cases and four compact reconstructions across four source versions: 93 successful requests, no retries. Three public browser evidence journeys then reconstructed the same care-home package exactly: 35 records, 50 relationships, insufficient evidence, no AI answer. Chrome and WebKit passed strict console checks; Firefox retained two hosting-cookie warnings. The 90-request browser run made no retries. The new release explicitly records historical browser journeys as not run.
The service-only Explorer merge exposed a publication gap: its nested guides
passed the selected PR jobs but Pages rejected missing top-level documentation
and changelog updates (35542056363). A separate correction updates both and
moves lockstep checking ahead of CI impact selection. It changes no runtime
bytes and retains the failed run. Publication of that correction remains open.
Both successor model controls passed event recognition and mechanical abstention.
The separately frozen protocol is bound to source commit
78a8beea97242d646eb9159860dea190ca5e2998, manifest
faefc7f42282c2f8dfdee119ff6d30f26b537b73b43b1b9ed133c3befbef4678.
Only the first substantive pair is authorised at this checkpoint. The original
five unsuccessful attempts retain their original outcomes.
Evidence and model checkpoint, 21 September shortly after midnight BST
The public full Reader passed a separate real Chrome journey against 270 unique immutable source files and all 21 application files. Concept filtering, the statutory graph, requested-version versus audit dates, the care-home heading and both unresolved SDA meanings were observed. The 16.6-second cumulative journey had no console or network errors. Its first harness reset-state assumption failed and remains retained. Both context packages remain insufficient and truncated. The public/local comparison preserves the different binding URL and derived context identities rather than asserting equal package IDs.
The successor model experiment is stable: nine attempts, seven parser-accepted responses, six mechanical passes including both empty-evidence controls. It retains 14 substantive claims and 21 citations, one defective source locator, a Claude formatter rejection and a Claude timeout. Three remaining Claude cases are explicitly held; no successful substantive pair is claimed. Model critique records scope and citation concerns separately from the mechanical results.
A further source review found that the care-home partner qualifications already exist in the captured corpus and authored semantic index. The small package removes them under byte pressure while retaining an authored summary. The task profile does not yet declare those qualifications required. The next bounded work separates a source-backed dependency correction from a generic allocation change: no DWP-specific routing in Explorer, no new acquisition claimed, and no rewriting earlier trials. Existing no-pressure context results must stay exact.
The presenter now has a Voice rehearsal sheet and a verified-evidence fallback. Actual Voice tool access and room audio remain untested. No account, microphone permission or audio setting was changed.
Editorial check: the eight affected overview, handover and verification documents
were checked with Explorer's British-English checker. Its six flags all concerned
judgment or judgments in references to judicial decisions. These intentionally
retain the UK legal spelling; none is an American-English prose substitution.
Publication and review checkpoint, 21 September at 00:20 BST
DWP checkpoint c19446c1 is pushed to PR 16. It includes the current service and
public Reader observations, frozen successor trials, Voice rehearsal and updated
beginner documentation. Private correspondence and unrelated research remain
outside the commit. Explorer PR 130 passed all checks, including the full browser
suite, and merged normally; canonical CI and Pages are being checked separately.
A separate offline critique checker verifies seven retained answers, 14 claims, 21 citation diagnostics and four additional selected records. It preserves the invalid source-locator pair and rejects evidence or authority changes. Eleven tamper controls include a check that changing a model opinion does not make it machine-validated semantic truth. No frozen experiment file was changed.
The next source and engine improvements have separate working trees. Qualification dependencies will be declared in DWP authoring; Explorer will allocate bounded context using those general dependencies. Recomputing missing declared dependencies after trimming deliberately corrects an existing diagnostic gap: an omitted dependency must remain visible even if its relationship was trimmed. Historical observations and model trials remain immutable.
Initial qualification source checkpoint, 21 September
The isolated source candidate adds eight required-support assertions and two explicit qualification profiles. It retains 901 records, now with 1,435 semantic assertions. All seven declared support pages already existed; no source text was acquired or edited. Producer tests and Explorer input validation pass. The combined Reader build preserves 20,044 records and now has 21,195 relationships. Its existing requirement labels apply in both directions; no new vocabulary or display fallback is necessary. Context allocation and publication are separate remaining checks, and no one has closed the 203 outstanding obligations.
Component qualification follow-up, 21 September
The bounded component review led to two narrower definitions and seven additional support relationships. The housing summary now describes the former-home treatment and Housing Benefit expenditure that may be payable. The temporary-residence summary keeps its no-partner opening. Five housing-cost pages and two temporary-residence pages are required support for their respective summaries. Staff 012 and 013 add the housing-cost closure; the permanent-care-home questions do not unconditionally require the temporary-residence branch. The broader no-partner and severe-disability overview support options remain explicit follow-up work.
The current semantic candidate has 901 records, 1,442 assertions, 15 support
dependencies and two qualification profiles. Each of Staff 012 and 013 declares
20 required identifiers and 17 paths; its five open obligation identifiers are
still absent evidence. All 203 obligations and original candidate identifiers
remain unchanged. Its snapshot is dwp-staff-semantics-fc2ad1545adcbeeb2243 and
index SHA-256 is 7ffc9d00e71fef6aed5531373510df82998123e89fdb28091cfb82384adf2876.
The regenerated combined Reader retains 20,044 records, 513 source PDFs and
19,090 source pages, with 21,202 relationships. Its snapshot is
dwp-combined-9b85571294b473b9fe6e. The 20 selected statutory units remain separate
from the manual counts. Current producer and Reader outputs are reviewable
engineering artefacts; the earlier eight-edge checkpoint above retains its
original counts. Joint context allocation, exact-version publication and answer
quality remain separate checks. No frozen source, model-trial package or browser
observation was rewritten.
Verification for this follow-up: all 31 staff producer controls, 20 combined
Reader controls and 15 backlog controls pass. The Reader checks cover all
15 requirement edges, their four source-concept groups, both navigation labels,
provenance and unreviewed authority. The backlog projection matches its register.
Both producer --check commands pass; the combined check reproduces all 4,758
declared output files exactly. These offline checks do not establish a later
browser or AI-answer result.
Joint qualification verification and ledger update, 21 September
- Root froze the final DWP candidate as
7f9feb9634e3d94004853b838462aca132c505a5and the independently reviewed Explorer allocator asc4f2de0a99b7bc2f8b8c8a06a3c715fb56b66d8e. The source has 901 records, 1,442 assertions and 15 explicit support relationships; all 203 obligations are unchanged. The combined projection has 20,044 records and 21,202 edges. - The semantic agent ran the four-cell comparison: two immutable sources, two archived engines, two budgets and 40 cases produce 320 assemblies. Actual run and exact deterministic replay pass. The Explorer agent independently reviewed archive admission and the authored-path census; all 14 offline controls pass. A review finding was corrected before execution: omitted output requirements cannot shrink the expected path denominator.
- The final pair retains 177 of 177 candidate occurrences and 433 of 433 activated declared paths at both budgets. Staff 012 and 013 retain all seven household pages, with 27 records/63 relationships at 256 KiB and 62/124 at 512 KiB. The latter packages expose a missing optional temporary-care-home dependency; no evidence sufficiency or specialist acceptance is claimed.
- The trade-off remains visible: on the old source at 512 KiB, allocation can displace six household pages that were not declared required. The new source declarations are necessary for the verified seven-page retention result.
- Root separately regenerated the current source-only evaluation: 12 to 177
candidate occurrences. At 64 KiB, Staff 012 correctly produces a 1,926-byte
metadata_budgetrefusal with zero records; the other seven budget controls retain non-empty evidence. All 40 main cases remain insufficient. - Documentation and the generated work-package ledger now distinguish completed bounded qualification retention, pending exact-version publication/public checks, larger unauthored component support sets and independent review. All earlier model attempts, source snapshots and public receipts are unchanged.
Joint receipt SHA-256: 45fb3ce7b768833eb6f94b6461f1531f56293de419665bd5adaf606d9bfdc3ac.
The evidence-path improvements do not retrospectively repair or regrade any
model answer. No new model or public HTTP calls were part of this comparison.
Continuing independent increments
Two follow-on implementations are running in isolated checkouts. The disability
addition work narrows the summaries, adds the two captured treated-receipt pages
and reviews their conditional dependencies. It is tracked as
DWP-BL-007.component-qualification-expansion; the current qualification release
and its fixed comparison are unchanged.
The service work adds explicit assembler identity and bounded historical replay
under DWP-BL-008.engine-replay. The existing progressive evidence-read milestone
retains its recorded completion; the aggregate item is now in progress because
this additional work has started. Public retained evidence resources have their
own pending package. Neither workstream has deployed a new service or rerun the
frozen model trials.
Qualification publication and disability follow-up, 21 September at 01:35 BST
Explorer PR 131 merged as 9cf67adb12cd39c6c008336de92929a08972a440.
Canonical run 35546895467 and Pages run 35546875345 both passed. A fresh
public Chrome observation at 01:23 BST bound the deployed application manifest
9fc8cb1bbf10e4e5182efd69d56f2b5ed39a2e6ecf942dce64357c4a529d1ce8 to the
qualification source 7f9feb9634e3d94004853b838462aca132c505a5. Six journeys
passed, with 270 exact corpus files, 294 responses and no console/network errors.
The care-home package has 62 records and 124 relationships; SDA retains both
unresolved meanings. This is a new observation, not a relabelled earlier receipt.
DWP PR 17 merged as 91b9907836c8a340d974dd958969f4f8cbb3a0c6 after both required
validation runs passed. A merge commit preserves the source/evaluation commits.
Canonical validation 35547851197 and the subsequent learning-site publication
are separate checks still under observation at this checkpoint.
The legal agent authored and the semantic agent independently reviewed the disability increment. Root retained both receipt qualifiers identified in review and the stricter existing housing wording control when integrating the two overlapping edits. All 46 semantic and 20 Reader controls pass. The semantic index has 903 records, 1,464 assertions, 98 selected pages and 29 support dependencies. All 203 obligation identifiers and statuses remain open; two source-closure labels are intentionally clearer.
Root rebuilt all 4,758 combined outputs and the current 40-case evaluation,
then verified exact regeneration/replay. The combined Reader has 20,044 records
and 21,224 relationships. The separate disability comparison
uses source 8ea4465a4cb5a867d82c87e635f2ef1d1df18d8a, two archived engines,
two budgets and all 40 question occurrences. Its 320 assemblies and exact replay
pass, with 14 admission/denominator controls. The current engine retains
497/497 paths at 512 KiB but 407/497 at 256 KiB. Candidate overlap remains 177/177
at both sizes; it cannot establish qualification coverage. All contexts remain
insufficient. The 64 KiB care-home request explicitly refuses without evidence.
The next partner-component source work is isolated and tracked separately. The service's explicit-engine replay implementation has passed independent runtime review after a shared-deadline correction; its final build/browser checks and deployment remain separate. A new direct-JSON subscription trial protocol has been proposed but no calls or new answer freeze have occurred. All earlier model failures, trial receipts and the live 0.5.0 service are preserved.
Publication checkpoint, 21 September at 01:50 BST
The qualification merge 91b9907836c8a340d974dd958969f4f8cbb3a0c6 now has
successful canonical validation 35547851197 and learning Pages publication
35548811691. The retained learning-site receipt still identifies the earlier
7815b17b… release; a successful hosting workflow is not a new HTTP observation.
Root reviewed and ran the new qualification public-receipt checker and all 14 controls. It verified 273 immutable source inputs, 21 application materials and 90 whole-evidence occurrences across the care-home and SDA packages. Every selected and returned required path is retained. Separately, 42 and 35 returned relationship rows point to 37 and 31 unselected targets with explicit budget omissions. The documentation now explains that those references do not supply the omitted evidence. All 11 original browser files remain unchanged.
The partner-component increment passed a separate source review: 38 staff and 14 household controls pass; the 903-record index has 1,482 assertions and 39 support dependencies. All 520 pre-existing evidence records and the identities, categories and statuses of all 203 obligations are preserved. Four conditional labels are intentionally clearer. Combined projection, bounded context retention and public delivery for that next increment remain separate work.
CI ordering correction
PR 18 run 35548995209 failed because the new receipt test executed before CI
fetched its immutable 7f9feb96… source. The local checkout already contained
that commit, so its 396 passing tests did not expose the shallow-checkout
dependency. The workflow now fetches that exact source before the unit suite;
the later comparison step reuses it. No test, receipt or source is weakened or
rewritten. The failed run remains visible; the corrected commit needs a fresh
required CI pass before merge.
The subsequent run 35549578095 reached all 396 tests and exposed a second
portability problem: six temporary-directory cases used macOS /private/tmp,
which does not exist on Ubuntu. The test now resolves Python's platform temporary
directory before constructing its isolated fixtures. The symlink checks and
frozen checker/receipt bytes are unchanged; all 14 focused controls and the
offline receipt check pass locally. A new Linux CI run remains required.
Partner retention checkpoint, 21 September shortly after 02:00 BST
The reviewed partner source is frozen at
7e5fdb9b906052914b307c17c0fd19feb2d008a7. Root rebuilt and checked all 4,758
combined outputs: 20,044 records, 21,242 relationships and snapshot
dwp-combined-a769a75d999ffd48afb0. All 402 Python controls, the 40-case current
staff evaluation and exact replay pass. All 203 obligations remain open.
The separate partner comparison runs 320 assemblies and exact replay. The legal agent independently checked 380 immutable input bindings, archived code, the 497-to-585 path denominator and eight exact focus-package reconstructions. Fourteen admission/census controls pass. At 512 KiB the current assembler retains all 585 declared path occurrences; at 256 KiB it retains 449. Candidate overlap is 177/177 at both budgets. Staff 012/013 retain 93/93 paths at the larger size and 25/93 at the smaller size. Missing qualifications remain visible.
The learning site at revision 91b99078… passed a new actual HTTP observation
at 01:53 BST: 123 matching responses, 1,777,711 bytes, 120 HTML pages and 2,382
internal links. The executed verifier and receipt are retained separately while
the next publication record is prepared. The disability candidate df352daa…
also passed six actual public Chrome journeys against application manifest
9fc8cb1b…, with 270 distinct corpus files and no console/network errors. Its
own source/receipt identity is separate from the earlier qualification check.
The service replay verifier passed independent review after stricter error-only response checks and pre-open special-file checks. All 77 service controls pass; no new public service call has occurred. The direct-JSON model runner has 24 offline controls and is under independent review; no provider has been called. Ignored-person discovery has produced a concrete source proposal, with its normal-residence distinction and remaining statutory/judgment gaps explicit.
Ignored-person coverage and capacity checkpoint, 21 September at 08:10 BST
The separately reviewed ignored-person source c44bc3a1… adds two conditional
concepts and eight already captured source pages. Root rebuilt and checked all
4,758 combined files: 20,046 records, 21,286 relationships and 19,912 resources.
All 411 Python controls and the 40-case current evaluation/replay pass. The
original 520 evidence records and 203 obligation objects remain unchanged.
The separate frozen comparison preserves 320 assemblies and exact replay. Eighteen admission/census controls pass. At 512 KiB the new source retains 177/177 candidate occurrences but only 647/765 declared path occurrences; at 256 KiB these fall to 171/177 and 393/765. The care-home cases lose one of seven household support pages even at 512 KiB. The byte census reproduces each measured package hash: Staff 012 contains only 45,608 JSON bytes of source evidence text, while requirements and relationships occupy 124,383 and 170,815 bytes. This identifies representational overhead for investigation; it does not justify silently dropping dependencies or declaring an answer complete.
The source modelling increment is separately recorded complete. Its capacity
regression has a named open work package. The Monday service and model trial
remain fixed to the partner source 723bcc5b…. Work was interrupted overnight
by usage limits; on resumption, source PR 19 passed and merged normally, and the
reviewed service passed all protected checks. New live observations and model
attempts are separate from these offline measurements.
Retained later public observations, 21 September
Root preserved the 01:53 BST learning-site check for exact source 91b99078…: 123 HTTP 200 responses, 120 HTML pages and 2,382 internal links. The separate 01:55 BST public Chrome observation of disability source df352daa… passed six journeys and retained the complete packages. Its offline admission checks 273 immutable source inputs, 71/71 returned required care-home paths and 12/12 returned required SDA paths; frontier relationships remain separately reported. Six new local checker-admission controls cover altered bytes, file bounds, symlinks and non-regular files. These observations do not attest the later partner source or the undeployed MCP successor. Historical files and .email.md remain untouched.
Versioned public delivery and retained examples, 21 September at 08:30 BST
Root deployed service 0.6.0 once, retaining the exact Worker, source and hosting identities. The first public SDK check delivered all 11 evidence cases but failed its comparison of different SDK envelopes. A reviewed verifier-only correction preserves complete tool rows and exact schema/trust checks. The fresh run passed 121 requests, including all nine compatible source/engine pairs, an empty control and exact historical replay. The failed run remains unchanged.
The reusable offline archive exporter passed independent review, eleven controls and a local seven-check Chrome journey. The DWP publisher passed 23 controls; its separate public byte verifier passed thirteen. Root exported three approved actual public packages into 225 files/2,766,291 bytes and admitted their exact Git-bound inputs. Public Pages and browser verification remain separate gates.
Both direct-v3 subscription controls completed within bounds but their parsers rejected undocumented-for-that-run metadata. The receipts preserve the failure categories and unknown tool census. No substantive call or retry followed. Installed-client schema inspection is informing a separately reviewed successor; the frozen v3 inputs remain unchanged. This is not a successful paired answer-quality or affordability result.
The first integrated 457-test run exposed one draft-only test that still required the protocol to be pending after the approved freeze. Its lifecycle assertion now requires the committed freeze while separately exercising refusal of a synthetic pending protocol. Frozen executable inputs and actual trial outputs are unchanged.
CI temporary-directory portability correction, 21 September
PR 21's eight retained direct-v3 observation controls failed on Linux because
the test fixture required the macOS-specific /private/tmp directory. The test
now uses Python's system temporary-directory default and resolves the resulting
path before applying the existing strict symlink checks. All eight controls pass
locally with both the ordinary environment and an explicit /tmp symlink on
macOS. Linux CI must confirm the candidate after publication; these local checks
do not claim a remote CI pass.
A scoped scan of test files changed in the previous 15 commits found no other
hard-coded platform temporary-directory requirement. The other guarded Python
fixtures already resolve tempfile.gettempdir(), while the browser controls use
Node's tmpdir(). Literal private paths in older negative tests are synthetic
rejection/redaction inputs, not directories to create. Only the fixture and
these documentation notes changed; frozen executable inputs and recorded
provider outcomes remain untouched.
The next push-run failure occurred during fixture cleanup: a background Git pack directory changed while Python removed a synthetic repository. Fixture repositories now disable automatic garbage collection and maintenance locally; real repository configuration and publication validation are unchanged. The failed run remains visible in CI.
Paired direct responses, 21 September
The separately reviewed v4 freeze preserved the complete v3 packages, prompt and answer schema. Its parser adds documented installed-client metadata recognition, bounded structural diagnostics and strict numeric usage fields. Thirty-nine controls and independent review passed before any provider call.
Both actual empty controls passed and abstained. The root then made exactly one Staff 012 call per subscription client: both passed mechanical checks with complete event census and zero observed tools. The two answers contain six claims and seven exact citations. Both distinguish the whole award from additional amounts and decline a whole-award conclusion. Independent agent claim review is separate; specialist acceptance, model identity parity, comparative accuracy and affordability remain unestablished. No trial was retried or earlier failure overwritten.
The independent agent review binds four attempts and 21 selected qualification records. It finds the six main scoped statements traceable, but records omitted treated-receipt/transitional exceptions, an overstated gap and further citation needs in one response. The answers remain unchanged and BL010 stays open for qualification repair and human assessment. The integrated local suite passes all 507 Python controls; public and protected-main checks remain separate gates.
Service 0.6.1 and native client checks, 21 September 2026
The question-schema patch was published as Sites version 11 at 11:52:32 BST.
The reviewed runtime 1420c316… and merged Explorer commit 68743984… share
complete Git tree 67eb52f7…; the original build identity is preserved. An
initial archive save was rejected before a version was created because of the
entrypoint layout. The corrected final archive used unchanged Worker bytes.
The 0.6.1 release record
links the exact archive, storage, runtime and publication identities.
The new actual SDK run passed 11 cases in 121 requests, receiving 10,322,722 bytes from 11:53:33 to 11:55:35 BST. It includes all nine allowed source/engine pairs, the historical care-home package and an empty control, without retries, model calls or full-package tool calls. Every complete package hash matches the preserved 0.6.0 run. The larger Staff 012 package remains insufficient.
After refresh, ChatGPT settings advertised the corrected pattern, three tools,
five sources and two engines; permissions were unchanged. The same previously
rejected native unknown-term control then succeeded. The exact Staff 012 native
smoke check also returned, but its 16 KiB budget retained no records and reported
an explicit byte-budget omission. It does not establish useful answerability.
The existing Codex task still exposes only full ask_okf; compact-client and
nested Data Agent access remain unproved. BL008.client-connection stays in
progress, and no new AI-answer or Voice acceptance is claimed.
The status checker was also exercised against real drift: new receipts with
the old 0.6.0 selection failed with “A newer successful deployment is not
represented”. Pinning the receipts at immutable commit a13a291f… produced the
0.6.1 status and passed check mode plus 19 controls. The before/after client
records and all older SDK failures remain separately retained.
Abroad semantic audit and reusable question diagnostics, 21 September
Root traced the reported five-page result to the 19 September pre-fix observation, then ran ten offline packages against exact source and engine identities. The published 723 source resolves the exact abroad question at 32 KiB, retaining two records and one relationship; overseas paraphrases remain unresolved. This is a local replay, not a new public or model observation.
Separate agents own the additive DWP international graph and its source/path tests, the reusable Explorer question-scaffolding classifier, and independent whole-passage review. Root owns integration, backlog, documentation and release boundaries. The existing four international IDs are reused; no frozen source, engine, receipt, trial or private email is rewritten. The audit names cross-benefit and graph-budget work still outstanding. The Data Analytics guide proposes a controlled semantic-proposal exercise; it is not a completed new model trial.
The bounded source repair passed independent agent review, 67 semantic tests, 21 combined Reader tests and the current 40-case replay. Its separate comparison retains 24 complete packages and 10 negative controls. Four Pension Credit cases retain 7/7 whole pages and 14/14 required paths at 512 KiB; the general 32 KiB result regresses to one concept and no source evidence. The regression is explicit in the retained report. Public-service admission and generic capacity repair remain separate; all 203 obligations stay open. Explorer's shared-classifier change passes 607 tests and has its own reviewed PR; frozen MCP engines are not rewritten.
At the owner's request, a separate ChatGPT Work task returned a frozen-file
semantic review and proposal summary for source 723bcc5b…. Its reported
workflow used local Git and file/PDF inspection after raw web reads failed;
it reported no callable Ask OKF tools or separate nested Data Agent call. The
local task inspected the returned messages and proposal summary, not the two
complete cloud artefacts. Their reported hashes remain unverified locally.
Novel proposals await full artefact import and independent source review; the
method guide records this boundary rather than
claiming native MCP acceptance or an accuracy benchmark.
Case-level wider semantic audit, 21 September 2026
A read-only audit inspected all 40 retained current after packages (39 distinct
questions) and verified their decoded hashes against the evaluation at source
commit c203a4bd621e57c99273b3933df0207e101c5a85. The packages use the pinned
c4f2de0a… engine, a 512 KiB limit and semantic index SHA-256
92a8871b8f1f2e51f1feace0b1f57c67dfd0ddcb0434cf4574795e5942e95fd6.
No assembly rerun, model call or public-service test was made for this audit.
The wider-work section now names the remaining work: 27/38 ADM-containing packages have no ADM relationship path (27/40 overall); 24/40 activate multiple profiles; and 23/40 retain unresolved tokens, including both domain constraints and ordinary wording. It distinguishes duplicated and broad profile triggers from missing evidence. Only cases 012/013 lose declared required paths: 118/776 path occurrences across the 40 packages, or 62/463 after deduplicating identical paths within each package. Separately, 15/40 have 86 support-dependency diagnostics, all pointing to records present in the index. Existing backlog packages cover these findings. All 40 packages remain insufficient and all 203 obligations stay open. These checks describe recorded selection and diagnostics, not specialist acceptance or a new legal answer-quality result.
The local browser review of the committed learning website exposed two stale learning-path descriptions: the returned Data review still read as unrun, and the dated 0.6.0 baseline was called current. Both are corrected. Mutable service status now points to the generated, receipt-checked record; historical counts retain their own dated observation. No service deployment was performed.