OKF-DWP

Independent experimental publication. Not an official DWP service, benefits advice or an entitlement calculator.

View this version’s Markdown source

On this page

Logical-unit implementation work log

Source-led manual structure and discovery cards, 22 September 2026

The owner authorised a complete DMG/ADM source-led build and an explicit active goal. The process plan records the scope and gates. The independent DWP worktree starts at 908adc45; the Explorer adapter worktree starts at ea485af6. Existing releases, private files and source captures remain untouched. The old Monday heartbeat is not restarted.

Root owns source structure and integration. The manual-guide agent owns exact source-supported reading conventions and the document-role census. The coverage agent independently freezes four source-bound acceptance cases and their doubled eight-case set before observations. The Explorer agent owns a scalable generic discovery-card and relationship adapter, keeping older corpus versions stable. Publication will have one owner while independent work continues separately.

The starting defect is broader than missing question profiles: inherited headings can be wrong, large literal-text candidates dominate ranking, and small delivery budgets omit evidence. Card descriptions remain distinct from domain concepts, full source passages and necessary legal support. No new result or improved answerability is claimed by recording this plan.

Started 22 September 2026 following the owner's explicit unattended implementation request. The earlier Monday deadline and paused heartbeat remain historical.

Source-led implementation checkpoint

The new feature branch has a source-supported manual guide and a separate bounded PDF-structure observation layer. All 513 PDFs were inspected locally; the raw observation history and a corrected bounded parser replay are retained. One observation reached its output bound and remains an explicit fallback.

The first complete candidate unit build accounts for 19,090 pages and 35,143,443 source bytes, with 49,491 non-empty units and all 75 authored records preserved. These are development build counts, not published-release acceptance. The initial context projection contains one separate discovery card per unit, 331 existing concepts and 20,632 graph assertions; 26,472 reference observations remain unresolved. Full Reader integration and the fixed forty-question comparison follow before any default changes.

The final bounded structural experiment passed four cases and then eight. Earlier failed attempts and additional independent generalisation failures are retained. Repeated-heading ambiguity, notice boundaries, fragment-local reference offsets and raw-tree/sidecar binding received regression controls. No legal, specialist or model-answer acceptance is inferred from these results.

Delivery and source-boundary corrections, 23 September 2026

The first full forty-question comparison is retained at context-probe/runs/attempt-01. It failed acceptance: at 32 KiB the new projection retained no evidence for any of the forty staff occurrences; at 512 KiB it retained more source units but only 40 of 47 previously declared paths. Eager discovery reads consumed the shared file allowance before concept routes, and full card/incident metadata consumed payload space. This is a delivery regression, not an answer-quality gain. The original engine and corpus remain recoverable at the recorded commits. The repair gives resolved-concept routes priority and uses exact hash-bound metadata references. The full comparison must be rerun before adoption.

A separate corpus-wide role census found reserved-number ranges and short illustrative notices carrying their labels into later appendices. Four exact source regressions and independent visual checks now support explicitly ended regions. Unclassified following material stays visible as unresolved source; it is not discarded or renamed a rule. Memo numbering remains a separate candidate observation because PDF paragraph tags alone also label nested points.

After these parser and literal-reference changes, the new candidate-initial-06 gate passed 4/4, followed by candidate-expanded-05 at 8/8. These are the same fixed development cases, with the new helper modules included in their implementation bindings. Earlier runs remain unchanged.

The runtime comparison retains the same source and authored semantics. New Pension Credit and household selection proposals are prepared separately so their gains or losses cannot be confused with a ranking or delivery repair.

The second runtime comparison completed all 164 assemblies with no network/model calls. At 512 KiB the new engine restores 47/47 declared paths and reduces median discovery diagnostics from 51,339 to 16,447.5 bytes. The 82 legacy-projection packages are byte-identical. At 32 KiB all forty new-corpus contexts still contain no evidence; every staff result remains insufficient. This repairs a measured regression without making an answer-quality claim.

Further source reading found a wrapped range endpoint being misread as paragraph 77164. A generic continuation check preserves the preceding paragraph's note and citation. Source regression, before/after identities and the concurrent fixture registration boundary are retained in the auxiliary review. Fresh gates candidate-initial-07 and candidate-expanded-06 pass 4/4 then 8/8. The corrected catalogue has 52,841 units, including all 75 unchanged authored units, and exactly 35,143,443 original source bytes. More unresolved fragments are now visible because erroneous reserved/notice labels no longer conceal the following text.

The unit-count reduction exposed a stale final generated shard. The failed post-write admission and exact surplus bytes are retained in build-history. Only that verified generated surplus was removed; deterministic producer replay then passed. No frozen source or earlier logical/page projection changed.

Starting state and ownership

DWP main: 82bd0a18e9941752798e5f4446eaa8bfb8b46936. Explorer main: 74e29f551776b27b08bfe3bacd774282f0408ffd. Both changes use isolated codex/logical-evidence-units branches.

Owner Exclusive implementation area Initial state
Producer agent DWP segmentation producer, tests and logical-units/ In progress
Source agent domain-profile/logical-units/ and independent source fixtures In progress
Explorer agent Reusable unit/corpus v2 support, dependency loading, UI and tests In progress
Integrator DWP logical runtime/Reader, evaluation, contracts, documentation, backlog and reviewed PRs In progress

The source/page and earlier bundle projections remain unchanged. No new source acquisition or model-provider calls are required to start. Private .email.md and unrelated research remain outside this change. Root Git checks found no competing feature writer; existing unrelated PRs are preserved.

Decisions

Implementation checkpoints, failures and measured outcomes will be appended as they occur. An implementation package is not complete merely because its design is recorded here.

Source and implementation checkpoint

Retained integration failures and corrections

  1. The first generated catalogue became stale while the source reviewer corrected fixture keys and spans. The integration build refused the mismatch. Final source input was rebuilt and frozen before the comparison resumed.
  2. The first allocation failure exposed a context authority mismatch: model-derived assertions had used the publication vocabulary class editorial. The producer now uses the existing context class model-assisted. The engine's governance check was preserved.
  3. A natural temporary-care phrase did not resolve its existing concept. An explicit source-backed alias addition is being tested; the corpus does not silently infer permanent residence from generic care-home wording.
  4. Peer review required explicit proposal-specific scope, cross-document references, bounded decompression and source recomputation before admission. These are implemented; final regression and publication results follow below.

No failed run has been represented as successful. No legal answer, specialist acceptance, model-quality improvement or new public service default is claimed.

  1. Actual browser loading initially refused the endpoint index: 100,507 entries exceeded the existing 100,000-entry safety bound because every unit duplicated its PDF resource. The Reader now shares one resource per source PDF, preserving unit-specific page spans. The consumer bound was not raised. Superseded unpublished generated resources were preserved in the local temporary archive; none belonged to a frozen release. Empty-source PDFs remain in the catalogue.
  2. Peer review added source-document owner consistency and guarded temporary-care alias authoring. All 11 context/Reader tests and 12 source tests then passed. The new resource-census regression is being added before final verification.

Final local checks and browser observation

Public Reader verification and CI correction

Additive closure and compact-delivery increment, 22 September 2026

Learning-release continuation and CI capacity

The subsequent learning release merged as bd94537742aeab0dcd016efaf04f898ef84c7f8e. Its reviewed-head checks passed, but merged run 35758898949 was cancelled by the runner's 30-minute limit; the annotation states that the maximum execution time was exceeded. The remaining checks were skipped and the dependent Pages publication did not run. A passing PR is therefore not recorded as a successful publication of the exact merge.

Increase the bounded aggregate job allowance to 45 minutes, retaining every source, semantic, replay, private-input and publication check. This adds timing headroom for the expanded corpus and learning checks; it neither changes source bytes nor weakens an evidence assertion. Keep the cancelled run available. The repair's PR handover records its exact checks, merge and subsequent Pages verification separately; changing this limit alone does not claim publication.

Separate discovery probes — 22 September 2026

Learning-site reliability correction — 22 September 2026

The preceding 548557f2 publication passed exact merged CI, automatic Pages, all 377 public-byte checks and its three retained evidence examples. Its later learning-page check exposed two defects: the focused skip link moved navigation before pointer activation, and aligned table cells produced 16 CSP errors. Those failures remain recorded; successful byte verification is not complete browser acceptance.

The independently reviewed repair is incorporated into the research follow-on PR. It keeps the focused skip link outside layout and converts supported table alignment to stylesheet classes, retaining the original CSP and inert source HTML. The regression guide records 26 renderer tests, eight browser checks at desktop/mobile widths and separate negative controls for both original defects. Fresh exact-head CI and public replay are required after integration. No source, probe result, profile, ranking or answerability status changed.

Manual-led complete-corpus increment, 23 September 2026

Expanded source selections and original-case check — 23 September 2026

23 September 2026: budget-limited Demo 1 freeze

The Demo 1 follow-up initially failed in the legacy job because its shallow checkout lacked the frozen source commit required by the five new tests. The structured job's fetch does not populate another job's Git objects. An isolated shallow clone reproduced five failures; the same pinned fetch made all five pass. The workflow setup is repaired without changing source, protocol, answers or acceptance criteria. The failure and reproduction remain under validation/demo-1-freeze/ci-shallow-checkout/.

Bounded Pension Credit capital candidate, 23 September 2026

The capital repair report records two exact source-boundary corrections, ten composite groups and a retained 52-package offline comparison. At 512 KiB the candidate retains 11/11 declared groups, compared with 9/11 for a newly authored topic-routing baseline; the apparent difference is six whitespace bytes. Both arms retain Staff 006 paragraph 84911 in full, and all packages remain insufficient. At 32 KiB neither arm delivers source evidence. The 11 external dependencies, compact delivery, specialist acceptance and production promotion remain open. No model calls or new source acquisition were made.

Capital pilot review follow-up, 23 September 2026

Retained all three fixed-question attempts. Corrected authored-summary provenance and reserved paragraph labels after agent review; the final 52-assembly replay retains the same substantive coverage and 32 KiB failure. The ten groups are a separately versioned experiment, not a production migration. See the repair report and reproduction commands.