Logical-unit implementation work log
Source-led manual structure and discovery cards, 22 September 2026
The owner authorised a complete DMG/ADM source-led build and an explicit active
goal. The process plan records the scope and gates.
The independent DWP worktree starts at 908adc45; the Explorer adapter worktree
starts at ea485af6. Existing releases, private files and source captures remain
untouched. The old Monday heartbeat is not restarted.
Root owns source structure and integration. The manual-guide agent owns exact source-supported reading conventions and the document-role census. The coverage agent independently freezes four source-bound acceptance cases and their doubled eight-case set before observations. The Explorer agent owns a scalable generic discovery-card and relationship adapter, keeping older corpus versions stable. Publication will have one owner while independent work continues separately.
The starting defect is broader than missing question profiles: inherited headings can be wrong, large literal-text candidates dominate ranking, and small delivery budgets omit evidence. Card descriptions remain distinct from domain concepts, full source passages and necessary legal support. No new result or improved answerability is claimed by recording this plan.
Started 22 September 2026 following the owner's explicit unattended implementation request. The earlier Monday deadline and paused heartbeat remain historical.
Source-led implementation checkpoint
The new feature branch has a source-supported manual guide and a separate bounded PDF-structure observation layer. All 513 PDFs were inspected locally; the raw observation history and a corrected bounded parser replay are retained. One observation reached its output bound and remains an explicit fallback.
The first complete candidate unit build accounts for 19,090 pages and 35,143,443 source bytes, with 49,491 non-empty units and all 75 authored records preserved. These are development build counts, not published-release acceptance. The initial context projection contains one separate discovery card per unit, 331 existing concepts and 20,632 graph assertions; 26,472 reference observations remain unresolved. Full Reader integration and the fixed forty-question comparison follow before any default changes.
The final bounded structural experiment passed four cases and then eight. Earlier failed attempts and additional independent generalisation failures are retained. Repeated-heading ambiguity, notice boundaries, fragment-local reference offsets and raw-tree/sidecar binding received regression controls. No legal, specialist or model-answer acceptance is inferred from these results.
Delivery and source-boundary corrections, 23 September 2026
The first full forty-question comparison is retained at
context-probe/runs/attempt-01.
It failed acceptance: at 32 KiB the new projection retained no evidence for any
of the forty staff occurrences; at 512 KiB it retained more source units but
only 40 of 47 previously declared paths. Eager discovery reads consumed the
shared file allowance before concept routes, and full card/incident metadata
consumed payload space. This is a delivery regression, not an answer-quality
gain. The original engine and corpus remain recoverable at the recorded commits.
The repair gives resolved-concept routes priority and uses exact hash-bound
metadata references. The full comparison must be rerun before adoption.
A separate corpus-wide role census found reserved-number ranges and short illustrative notices carrying their labels into later appendices. Four exact source regressions and independent visual checks now support explicitly ended regions. Unclassified following material stays visible as unresolved source; it is not discarded or renamed a rule. Memo numbering remains a separate candidate observation because PDF paragraph tags alone also label nested points.
After these parser and literal-reference changes, the new
candidate-initial-06
gate passed 4/4, followed by
candidate-expanded-05
at 8/8. These are the same fixed development cases, with the new helper modules
included in their implementation bindings. Earlier runs remain unchanged.
The runtime comparison retains the same source and authored semantics. New Pension Credit and household selection proposals are prepared separately so their gains or losses cannot be confused with a ranking or delivery repair.
The second runtime comparison completed all 164 assemblies with no network/model calls. At 512 KiB the new engine restores 47/47 declared paths and reduces median discovery diagnostics from 51,339 to 16,447.5 bytes. The 82 legacy-projection packages are byte-identical. At 32 KiB all forty new-corpus contexts still contain no evidence; every staff result remains insufficient. This repairs a measured regression without making an answer-quality claim.
Further source reading found a wrapped range endpoint being misread as paragraph
77164. A generic continuation check preserves the preceding paragraph's note and
citation. Source regression, before/after identities and the concurrent fixture
registration boundary are retained in the auxiliary review. Fresh gates
candidate-initial-07 and candidate-expanded-06 pass 4/4 then 8/8. The corrected
catalogue has 52,841 units, including all 75 unchanged authored units, and exactly
35,143,443 original source bytes. More unresolved fragments are now visible
because erroneous reserved/notice labels no longer conceal the following text.
The unit-count reduction exposed a stale final generated shard. The failed
post-write admission and exact surplus bytes are retained in
build-history.
Only that verified generated surplus was removed; deterministic producer replay
then passed. No frozen source or earlier logical/page projection changed.
Starting state and ownership
DWP main: 82bd0a18e9941752798e5f4446eaa8bfb8b46936.
Explorer main: 74e29f551776b27b08bfe3bacd774282f0408ffd.
Both changes use isolated codex/logical-evidence-units branches.
| Owner | Exclusive implementation area | Initial state |
|---|---|---|
| Producer agent | DWP segmentation producer, tests and logical-units/ |
In progress |
| Source agent | domain-profile/logical-units/ and independent source fixtures |
In progress |
| Explorer agent | Reusable unit/corpus v2 support, dependency loading, UI and tests | In progress |
| Integrator | DWP logical runtime/Reader, evaluation, contracts, documentation, backlog and reviewed PRs | In progress |
The source/page and earlier bundle projections remain unchanged. No new source
acquisition or model-provider calls are required to start. Private .email.md
and unrelated research remain outside this change. Root Git checks found no
competing feature writer; existing unrelated PRs are preserved.
Decisions
- Add new unit identities and projections instead of redefining old page IDs.
- Keep source segmentation status distinct from semantic and specialist review.
- Preserve old corpus v1 validation; v2 separates physical-page and unit counts.
- Keep aggregate literal hashes separate from exact fragment/span hashes.
- Load only bounded, hash-bound declared dependencies. Requirements do not become hidden retrieval seeds.
- Publish a new service default only after meaningful source, retrieval, qualification, delivery and live checks justify it.
Implementation checkpoints, failures and measured outcomes will be appended as they occur. An implementation package is not complete merely because its design is recorded here.
Source and implementation checkpoint
- Full source producer: 513 documents, 19,090 pages, 49,680 unit records and 35,143,443 extracted bytes accounted exactly; no unassigned source bytes.
- The 46 authored excerpts cover 42,864 source bytes in six documents, including 18 cross-page excerpts. The other 49,634 units remain uncertain machine candidates. These are different review denominators.
- Source authoring includes 28 scoped relationship proposals, two new neutral concepts, five task profiles and eight explicitly unresolved references.
- Producer: 19 controls pass. Independent source fixtures: 12 controls pass, including 43 qualification/example checks and a rehashed missing-footnote rejection. Every unit passes the new consumer shape and fragment-hash checks.
- The new corpus and indexed Reader have been generated from the same units; their final consumer and publication checks are still in progress.
Retained integration failures and corrections
- The first generated catalogue became stale while the source reviewer corrected fixture keys and spans. The integration build refused the mismatch. Final source input was rebuilt and frozen before the comparison resumed.
- The first allocation failure
exposed a context authority mismatch: model-derived assertions had used the
publication vocabulary class
editorial. The producer now uses the existing context classmodel-assisted. The engine's governance check was preserved. - A natural temporary-care phrase did not resolve its existing concept. An explicit source-backed alias addition is being tested; the corpus does not silently infer permanent residence from generic care-home wording.
- Peer review required explicit proposal-specific scope, cross-document references, bounded decompression and source recomputation before admission. These are implemented; final regression and publication results follow below.
No failed run has been represented as successful. No legal answer, specialist acceptance, model-quality improvement or new public service default is claimed.
- Actual browser loading initially refused the endpoint index: 100,507 entries exceeded the existing 100,000-entry safety bound because every unit duplicated its PDF resource. The Reader now shares one resource per source PDF, preserving unit-specific page spans. The consumer bound was not raised. Superseded unpublished generated resources were preserved in the local temporary archive; none belonged to a frozen release. Empty-source PDFs remain in the catalogue.
- Peer review added source-document owner consistency and guarded temporary-care alias authoring. All 11 context/Reader tests and 12 source tests then passed. The new resource-census regression is being added before final verification.
Final local checks and browser observation
-
Completed the 19 producer, 12 independent source and 14 Reader/context controls. Reader census: 50,524 records, 513 unique resources, 51,884 endpoint labels, 69,163 search tokens and 2,734,358 uncapped postings.
-
Actual browser Search found a second inherited bound mismatch: the producer declared total records (50,524) as the maximum postings for a token. It now reports the actual 49,635 maximum and rejects future overflow. No postings were removed and no consumer bound was weakened.
-
Independent review strengthened the evaluation itself: exact four-module allowlist and immutable Git blobs before execution; actual serialised bytes, node/edge/path bounds, fragment integrity, bounded archive decoding and exclusive observation-directory creation.
-
The final 184-assembly comparison and seven controls pass; 20 focused packages and the exact evaluator are retained. Preliminary observations remain in
preliminary-run/andpre-search-run/; they are not the final build receipt. -
Browser UI and actual read-only WebMCP calls returned the same PC-abroad context ID. The gateway unit displays both complete examples across PDF pages 8–9, exact spans and hashes; the UI shows source references and support dependencies. Insufficiency, unresolved references, candidate truncation and byte omissions remain visible. No external model answered this test.
-
At 32 KiB every focused unit case retains zero evidence. At 512 KiB all 24 new focused paths are retained; this does not establish complete legal scope or migrate all older profiles. These limits are explicit backlog packages.
-
.email.mdremains ignored and untracked. Frozen source, pilot, full-DMG and combined projections remain unchanged. Exact public CI and deployment verification are recorded in the DWP PR handover, separately from the checked-in local observation. -
Full DWP Python regression suite: 578 tests passed, including expected negative CLI controls. Canonical repository contracts, private-input check and exact 4,900-file logical-context rebuild pass. The local browser observation records separate source/app bindings; it is not a public deployment receipt.
-
Final independent editorial review corrected three claims before publication: machine candidates are not all complete passages; the care-home profile keeps temporary/permanent alternatives; the ADM memo is an unresolved reference, not a resolved graph destination. These corrections do not change the source or evaluated projection.
-
Delivery PRs: DWP 28 and Explorer 140. The pinned DWP Reader source is
adfa7137d3b24033d7265123c739a9ffcfb06584; later documentation corrections preserve its unit and context bytes. Exact merge/deployment observations belong to the PR handover, not an inferred deployment status from this source document. -
Updated the portable methodology and retrospective: establish source-bound logical units before semantic generation; keep boundary accuracy, dependency completeness, source census and answer quality as separate checks.
Public Reader verification and CI correction
- Explorer PR 140 merged
as
69d38b1c3c17939236f880e391746bf794623dda; all required head checks passed. Pages run 35741046500 rebuilt and publicly verified that exact merge. - The separate DWP public observation verified 50,524 Reader records, 513 sources, seven Search results for 077001, complete cross-page examples and directed support. The public UI JSON equals the actual WebMCP build result; explain returns the same context ID. The 440,273-byte package remains insufficient, with 18 records and eight edges. Exact public HTTP hashes and the Explorer deployment receipt are retained beside that observation. Old local warnings are identified separately.
- DWP CI 35740413305 caught unregistered work-package evidence: child entries referenced new files missing from their parent evidence lists. Registered those same existing files at the parent level; the validator was not weakened. Moved its cheap check before expensive corpus validation so future errors fail promptly. The failed run remains available; successful unit/source checks are not represented as an overall successful CI run.
Additive closure and compact-delivery increment, 22 September 2026
-
Isolated DWP branch
codex/semantic-compact-closurestarts at mergedbd945377; Explorer compact branch starts at mergedbf38d835. Learning sources, combined teaching descriptor, public assessor roster and private correspondence remain unchanged. -
Source reviewer owns dependency destinations; a separate reviewer owns the PIP transition slice and coverage ledger; Explorer agent owns generic compact delivery; root integrates producers, evaluation and documentation. Independent reviews found no blocking implementation issues; a duplicate-obligation-ID control was strengthened.
-
Added 29 excerpts across existing frozen sources, retaining 73,338 additional exact source bytes. The current 75 authored units and 49,551 uncertain candidates account for every source byte. Twelve references remain legally unresolved; target identification is a separate status.
-
Preserved the active 184-context baseline and engine bytes under
evaluation/logical-units/history/pre-closure-2026-09-22/before a fresh run. Retained source, earlier projections, evaluations and trial receipts are not relabelled. -
The initial regeneration correctly refused obsolete generated shard
records/0390.json.gzafter whole-unit grouping reduced the record census. Superseded generated files were hash-checked and preserved in a separate local temporary archive; historical committed bytes remain available. The exact new unit producer check then passed. -
One negative test initially expected the later document-owner error. The registered loader now rejects the deliberately altered reference earlier as an unknown disposition; the test recognises that fail-closed error without allowing the mutation.
-
Fresh evaluation passes 184 assemblies and 11 controls, with 20 exact focused archives, no network and no models. All results remain insufficient. Focused 512 KiB paths are 10/10, 9/9, 12/12, 2/2 and 4/4; all five focused 32 KiB inline packages still retain zero evidence after source expansion.
-
The coverage ledger records 8/40 staff occurrences activating a profile and 32/40 without one; 22/40 large staff contexts have no returned relationships. All 203 earlier obligations remain recorded. These are separate measures, not an answer-quality score.
-
Exact protected CI, merge and live publication checks follow independently; implementation and local results above are not a claim of a new remote-service default.
-
Independent review asked for positive assertions for Staff 026/033 and 032, beyond the existing five focused cases. Added exact profile activation, non-empty required paths, path retention and authority checks. Preserved the preceding run and coverage ledger under
pre-positive-controls-2026-09-22before replaying; the source and context identities did not change. -
The eight profile activations include broad Staff 029/036 DLA/PIP mentions. Their relationship-scope gap remains explicit. Activation is reported separately from intended task coverage.
-
Independent review found a ledger-only identity defect: dictionary field order let an authored short obligation ID overwrite its canonical identifier. Preserved the affected report under
evaluation/semantic-coverage/history/pre-canonical-obligation-ids-2026-09-22/, repaired the field order, retained both IDs and verified exact membership of all 203 unique canonical obligations in their earlier requirement sets. Eight coverage controls pass; the source and context packages were unaffected. -
Owner feedback challenged the limited 8/40 profile activation. A parallel diagnosis records the literal-text/alias gap and 16 heading-range mismatches in a bounded 26-candidate sample. The discovery-card experiment is a separate, unimplemented backlog package, not a claim that the wider retrieval problem is fixed. All 32 unprofiled cases return candidate evidence; that does not establish relevance or answer support.
Learning-release continuation and CI capacity
The subsequent learning release merged as bd94537742aeab0dcd016efaf04f898ef84c7f8e.
Its reviewed-head checks passed, but merged run 35758898949
was cancelled by the runner's 30-minute limit; the annotation states that the
maximum execution time was exceeded. The remaining checks were skipped and
the dependent Pages publication did not run. A passing PR is therefore not
recorded as a successful publication of the exact merge.
Increase the bounded aggregate job allowance to 45 minutes, retaining every source, semantic, replay, private-input and publication check. This adds timing headroom for the expanded corpus and learning checks; it neither changes source bytes nor weakens an evidence assertion. Keep the cancelled run available. The repair's PR handover records its exact checks, merge and subsequent Pages verification separately; changing this limit alone does not claim publication.
Separate discovery probes — 22 September 2026
- Root owns the isolated
codex/discovery-cards-pilotbranch. These diagnostics preserve source candidate7412d1d0, the approved engine2e557f4b, all 40 question occurrences and every existing source/profile/context output. - Alias protocol
f00d299awas frozen before implementation. Its reconstructed baseline agrees with the real engine on all 40 lexical candidate orders. Expansion recovers no earlier candidate location and loses two. Reject this shortcut for adoption. The initial unsupportedmax_edgescall and correction remain recorded; no validator was relaxed. - Length-aware protocol
e1654a4fwas frozen before implementation. BM25 with fixedk1=1.2,b=0.75recovers earlier candidate locations for four cases, loses none, and reduces median text in the first 16 candidates from 367,376 to 13,617 bytes. These are pre-graph ranking observations, not complete delivered evidence or correct answers. 626/640 selected occurrences retain unresolved machine boundaries; wrong-subject and historical matches remain. - The coverage-review agent independently verified both probes' bound inputs and recomputed every ranking. Reviews are retained beside the runners and observations. No source acquisition, provider calls or runtime changes were made. Structural and semantic cards remain separate unimplemented work.
- Exact protected publication of this research increment is a separate gate. Do not interrupt the preceding DWP source release's merged-CI/Pages verification.
Learning-site reliability correction — 22 September 2026
The preceding 548557f2 publication passed exact merged CI, automatic Pages,
all 377 public-byte checks and its three retained evidence examples. Its later
learning-page check exposed two defects: the focused skip link moved navigation
before pointer activation, and aligned table cells produced 16 CSP errors.
Those failures remain recorded; successful byte verification is not complete
browser acceptance.
The independently reviewed repair is incorporated into the research follow-on PR. It keeps the focused skip link outside layout and converts supported table alignment to stylesheet classes, retaining the original CSP and inert source HTML. The regression guide records 26 renderer tests, eight browser checks at desktop/mobile widths and separate negative controls for both original defects. Fresh exact-head CI and public replay are required after integration. No source, probe result, profile, ranking or answerability status changed.
Manual-led complete-corpus increment, 23 September 2026
-
Preserved the failed first discovery comparison. The second, using the exact same source and authored semantics, restores all 47 declared paths at 512 KiB; all 82 legacy packages are byte-identical. Median discovery diagnostic bytes fall from 51,339 to 16,447.5. All 40 small 32 KiB assemblies still retain no evidence, and every result remains insufficient. These are delivery results, not answer-accuracy results.
-
The final source-parser candidate accounts for all 513 documents, 19,090 pages and 35,143,443 extracted source bytes in 53,727 units. All 75 authored units retain their identities and bytes. Source-backed repairs keep reserved ranges, short notices, wrapped references, memo sections and appendices from being absorbed into unrelated paragraph units. Initial gate 08 passes 4/4, then expanded gate 07 passes 8/8; three extra memo/appendix controls and the earlier failed outputs remain separate. This is not perfect segmentation.
-
Add seven agent-source-read selection profiles and 55 conjunctively guarded edges. Add a separate location-only migration of the earlier 42 research candidate locations: 39 fully mapped, three with explicit whitespace gaps. Its 344 navigation routes comprise 312 restored routes and 32 inferred profile associations. All 40 original requirements and 203 obligations stay unchanged and open. Exact location overlap does not establish relevance, semantic equivalence, legal prerequisites or requirement support.
-
The first final-context build stopped before writing outputs because the Reader's postings limit did not admit the larger unit catalogue. Retain that failure, repair the generic producer/consumer contract, and run a separately bound full-question evaluation before adoption. Do not reuse the earlier runtime-only comparison as evidence for these new semantics.
-
The repaired Reader retains all 3,037,000 postings and their true frequencies. Two metadata tokens exceed the existing 50,000 completeness threshold; their exact counts are recorded and Explorer retains conservative uncertainty. No source record or posting is dropped. The next context build succeeds with 53,727 cards, 18,904 assertions and 54 separately grouped requirements.
-
Full trial 03 retains all 47 inherited paths and 139/181 additional source-read path occurrences at 512 KiB. Source-read profiles activate for 21/40 question occurrences; the separate legacy/location plane activates for all 40. All 203 original obligations remain open. One large-budget case returns no source evidence after metadata pressure and trimming, so adoption remains held. Frozen trial 03 and its independent report preserve this new allocation regression before a same-source runtime repair.
-
Trial 04 fixes the zero-source large package but regresses required paths: all 40 cases retain source, yet inherited paths fall to 46/47 and additional paths to 137/181. All 82 legacy packages remain byte-identical. The full result and independent analysis are retained; adoption stays held. A further reviewed runtime repair prioritises only paths actually observed from resolved seeds. A requirement cannot invent a seed or an edge. Its new 2,000-prefix work bound reports truncation and falls back to ordinary priority. Boundary uncertainty continues to make the package insufficient, while exact source can still be retained for review. Trial 05 holds the corpus, questions, protocol and ranking parameters fixed.
-
A separate original-acceptance audit finds that the source-led projection retains imprisonment concepts but omits the earlier custody requirements. Literal chapter references are not yet resolved by the paragraph-only reference producer. This migration gap is separate from the 40 staff cases; it must be restored and tested before claiming that all earlier hard cases work in the new projection. The earlier hospital case is an insufficiency control, not a previously completed legal answer.
Expanded source selections and original-case check — 23 September 2026
- The final expanded corpus accounts for all 53,727 units and now has 19,193 relationships and 82 separately grouped requirements. Global check 04 verifies every card/unit pair, source span, relationship and 1,355 input bindings.
- Trial 06 keeps the same engine, questions, ranking and budgets. At 512 KiB, all 40 staff occurrences retain source and relationships. Source-read activation rises to 37/40; 47 inherited, 181 earlier and 208 new declared source-selection path occurrences all survive. The three underspecified tasks and all 203 old obligations remain open. At 32 KiB no staff package retains source evidence.
- Location navigation remains a separate, less favourable measure: 469/765 restored and 34/87 inferred route incidences survive, compared with 490/765 and 35/87 in trial 05. The original page requirements are unchanged.
- Independent source review restores six original custody requirements, 53 narrowly selected units and four chapter routes. The separate first imprisonment assembly retains only 20/53 units and 43/89 paths at 512 KiB; all four complete chapter paths and the 12003-to-12002 dependency are lost. Retain this failure and diagnose its allocation boundary before claiming original-case acceptance. Hospital remains an insufficiency control.
- The Reader endpoint failure is repaired by indexing discovery aliases in their existing metadata channel rather than duplicating them as generic tags. No alias, source text, posting or conceptual facet was discarded and no limit was raised. The full Reader now loads 54,578 records and 19,193 relationships in the actual browser. Public acceptance remains separate.
- Explorer PR 144 merged through the normal protected squash route at
dc54fac4f98fa5bc9db38bf843854ce24a4c130d. The five tested runtime modules remain byte-identical to the separately pinned8a5b8d11engine.
23 September 2026: budget-limited Demo 1 freeze
- The supplied public brief exactly matches the forty existing question occurrences. The source candidate is in PR 33; its accepted staff replay reproduces all 164 assemblies without network or model calls.
- The owner authorised two paired questions, four new answer calls total. The frozen protocol and all four outcomes are retained. No calls remain. Twelve exact source quotations do not erase three format failures, ambiguous wording or the model confusing unmet requirements with absent profiles. No token-saving or specialist-acceptance claim is supported.
- Public Explorer loaded the pinned 53,727-unit corpus. Native in-app WebMCP returned two catalogue pages and one complete care-home source passage matching the human-visible context and frozen literal hash. Other clients remain separately unverified; the remote-service default was not changed.
- Local documentation rendered 193 pages; eight desktop/mobile journey checks passed. Exact merged CI and public Pages remain separate release gates.
- Separate imprisonment allocation and historical-amendment repairs remain parked on their own branches. The question ledger, claim review and freeze landing page are the demonstration handover; retained answers avoid rehearsal model calls.
The Demo 1 follow-up initially failed in the legacy job because its shallow
checkout lacked the frozen source commit required by the five new tests. The
structured job's fetch does not populate another job's Git objects. An isolated
shallow clone reproduced five failures; the same pinned fetch made all five pass.
The workflow setup is repaired without changing source, protocol, answers or
acceptance criteria. The failure and reproduction remain under
validation/demo-1-freeze/ci-shallow-checkout/.
Bounded Pension Credit capital candidate, 23 September 2026
The capital repair report records two exact source-boundary corrections, ten composite groups and a retained 52-package offline comparison. At 512 KiB the candidate retains 11/11 declared groups, compared with 9/11 for a newly authored topic-routing baseline; the apparent difference is six whitespace bytes. Both arms retain Staff 006 paragraph 84911 in full, and all packages remain insufficient. At 32 KiB neither arm delivers source evidence. The 11 external dependencies, compact delivery, specialist acceptance and production promotion remain open. No model calls or new source acquisition were made.
Capital pilot review follow-up, 23 September 2026
Retained all three fixed-question attempts. Corrected authored-summary provenance and reserved paragraph labels after agent review; the final 52-assembly replay retains the same substantive coverage and 32 KiB failure. The ten groups are a separately versioned experiment, not a production migration. See the repair report and reproduction commands.