OKF-DWP

Independent experimental publication. Not an official DWP service, benefits advice or an entitlement calculator.

View this version’s Markdown source

On this page

DWP guidance as an OKF+ bundle

Evidence workbench guide: inspect the 40 public staff questions, their selected source passages, relationships and explicit gaps. Read the 24 September handover for the verified public release, demonstration steps and remaining limitations. The blocked calculation inspection guide explains the opt-in, source-bound Pension Credit model and directional carer interactions. It does not calculate an award.

Demo 1 freeze plan for 23 September: tested source-led evidence, forty-question gaps and a four-call comparison cap.

An independent, unofficial experimental exemplar. Not an official DWP document, benefits advice or an entitlement calculator.

What changed: changelog · Current work log · Backlog and acceptance checks · How to repeat the method · What we learned · Monday handover and demonstration

This repository turns public Department for Work and Pensions (DWP) guidance into source-linked records, a YAML-LD semantic graph and an indexed OKF Explorer research candidate. It began with a Pension Credit pilot covering volumes 13 and 14, expanded to the full Decision makers’ guide (DMG), and now also preserves the separate Advice for decision making (ADM) manual. It demonstrates how a specialist or an AI can find evidence, inspect relationships and see what remains uncertain.

New to the project? Open the learning website or follow its Markdown source, or read how the web edition is published, then use the plain-English glossary when a benefit name or technical term appears. It explains Search, Ask OKF and AI answering through short tasks, including how to connect and why a browser can show HTTP 405.

For the semantic design, see the ontologies and namespaces actually used. For the separate model trials, see fixed-evidence answer review. The changelog records dated deliveries; the backlog keeps unfinished work explicit. Neither replaces the exact receipts below.

Verified Pension Credit semantic exemplar · Original meeting demonstration · Read the pilot bundle · Ten-minute meeting walkthrough · Discovery findings · AI interrogation guide · Public notice and rights

Current public release: versioned evidence and paired review

Use the latest recorded service publication status for the selected deployment and SDK observations. This generated page is the shared status reference for DWP and Explorer; it does not claim real-time health.

Dated observation: 21 September 2026

The public evidence service was recorded as 0.6.0, using the partner-qualified source 723bcc5b015ab38a026625c2148edbd784edf7c7. Its actual public SDK check reconstructed 11 evidence cases in 121 requests, including exact historical replay. The current care-home package has 55 records and 115 relationships and remains insufficient. Three retained examples let a person inspect small evidence parts and the same complete package; publication and public-browser checks have their own dated receipts.

The paired direct-v4 trial retains both clients' answers to the original care-home question and both empty controls. All four pass mechanical checks with zero observed tool calls. The separate agent critique records omitted exceptions and an overstated gap in one answer. Specialist acceptance and comparative accuracy are not established. Use the Monday handover for the current source distinctions, remaining work and ten-minute demonstration.

Earlier public 0.5.0 checkpoint: household evidence

The earlier public evidence service ran 0.5.0, using fixed DWP source 3ef0e786e9a18e76fa17c7d925ff509d6d6c9f84. The Monday handover gives a ten-minute route through the evidence and its limits.

The household evidence increment adds complete qualifying pages and selected dated statutory text. Its source-version evaluation retained 176 of 177 known candidate-page occurrences across the staff questions. All 40 packages remain insufficient, with 203 named obligations still open. This measures evidence discovery, not answer accuracy or specialist acceptance. The 20 selected statutory units are additional source extracts, not 20 complete Acts or a complete legal dependency set. They do not change the 513-PDF manual count below. Performance evidence, the new trial protocol and the Monday delivery log record separate engineering, delivery and answer-quality checks.

The actual SDK receipt passes seven full-package cases and four compact cases across four approved source versions. SDK means software development kit: here it is the external client used to compare service output with the shared evidence engine. The public browser observation passes the care-home evidence journey in Chrome, Firefox and WebKit. Chrome and WebKit also pass strict console checks; Firefox retains two hosting-cookie warnings. The care-home package contains 35 records and 50 relationships at a 262,144-byte budget and remains insufficient and truncated. No AI answer is generated. The separate historical browser journey suite was not run for 0.5.0; the earlier version's observations remain historical.

An earlier real public Chrome journey verified conceptual filtering, statutory text and graph links, source/audit dates, the care-home heading and unresolved SDA branches against 270 immutable corpus files. It had no console or network errors. Both assembled packages remained insufficient and truncated. Its cumulative 16.6-second run is one observation, not a general speed guarantee.

The household Reader checks also pass in Chrome, Firefox and WebKit against the exact local candidate. They verify all 20 statutory extracts, useful directed relationships, date distinctions, the care-home qualification and both unresolved SDA meanings. They do not establish public deployment or complete legal applicability.

Team handover and current work

The source-led manual guide explains the next additive pipeline: understand each manual's local conventions, propose coherent content units, then use compact discovery cards to find the underlying evidence. Its source census covers the same 513 captured DMG and ADM documents. PDF pages remain exact checking locations; headings declared inside a PDF are structural evidence, not proof of legal meaning. The guide records the candidate's tests, failed attempts and remaining gaps separately from earlier service releases.

The source-only projection currently contains 53,727 retrievable units across 19,090 PDF pages, including the 75 earlier author-declared units. The remaining boundaries are machine proposals. These counts describe source organisation; they are not counts of complete rules or answerable questions. The separately authored selection proposals add conditional routes to source passages. The location-migration ledger records how earlier page excerpts overlap those units while preserving all 40 staff requirements and 203 open obligations. Runtime acceptance and public adoption are separate.

The earlier logical-unit closure report records its frozen baseline: 75 authored boundaries, seven bounded profiles and 8/40 staff occurrences activating a profile. Its 184 contexts remain insufficient. The newer structured candidate adds 23 source-read profiles, bringing source-read activation to 37/40 staff occurrences in trial 06. At 512 KiB all 40 retain source passages and relationships; all 436 declared source-selection path occurrences survive. The separate location-navigation layer activates for all 40. These measures do not establish semantic closure or answerability: all results remain insufficient and the 32 KiB packages contain no source evidence. See the results and remaining work and source-led demonstration.

The bounded Pension Credit capital repair records an additive Chapter 84 candidate and its 52 offline assemblies. All packages remain insufficient; compact delivery, dependency closure and production promotion remain open.

The logical evidence unit work separates PDF page locations from the passages Ask OKF retrieves. It adds exact source spans, explicit supporting passages and a separate Reader/corpus projection. See its implementation log for tested progress and publication status; frozen page releases and service observations retain their original scope.

The additive staff increment provides projected personas and journeys, source-backed concepts and task requirements, legal reference reconciliation, a combined DMG and ADM Reader and a fresh source-listing comparison. Open the earlier staff-source Reader. Its public Chrome, Firefox and WebKit checks passed the declared navigation, evidence and keyboard journeys, verifying actual application and corpus bytes. These are reviewable research outputs. All supplied tasks retain explicit gaps; no specialist acceptance or complete benefits answer is claimed.

The later qualification-source public Chrome observation checks the deployed application against source 7f9feb96…. Six journeys pass, including the seven household support pages, conceptual filtering and the statutory graph. The retained care-home package has 62 records and 124 relationships. Its checker verifies 273 immutable input files and distinguishes selected evidence from graph references whose target was omitted by the budget. This observation predates the disability increment described below.

The Monday handover links the current demonstration and remaining semantic work. The 20 September team handover preserves the earlier checkpoint. The delivery and acceptance ledger separates implementation from independent review, so a review gate cannot hide unfinished work. The bounded staff concepts/relationships (BL005) and task profiles (BL007) are delivered. Broader domain expansion and evidence closure remain explicit implementation packages; mention classifications do not complete those packages.

The preceding verified qualification candidate declares 15 explicit support relationships: the initial eight household dependencies, five housing-cost dependencies and two temporary-residence dependencies. Seven captured pages support the complete household summary, including the distinction between one and both partners living in a care home. Staff 012 and 013 require the household and housing-cost paths. The temporary summary keeps its own support requirements without being made mandatory for every permanent-care-home question. That source index has 901 records and 1,442 assertions. Source text and all 203 open obligations are preserved. The joint source/engine comparison now passes all 320 assemblies and exact replay, using DWP 7f9feb9634e3d94004853b838462aca132c505a5 and Explorer c4f2de0a99b7bc2f8b8c8a06a3c715fb56b66d8e. At both 256 KiB and 512 KiB it retains 177 of 177 candidate occurrences and all 433 declared path occurrences. Staff 012 and 013 retain all seven household support pages. At 256 KiB each package has 27 records and 63 relationships; at 512 KiB it has 62 records and 124 relationships, with an explicitly missing optional temporary-care-home dependency. All contexts remain insufficient and truncated.

A smaller 64 KiB Staff 012 request legitimately returns a 1,926-byte metadata_budget refusal with no records: even its interpretation and obligation metadata cannot fit. Its public service update remains separate from the website and offline checks; the service links above retain their verified source versions. No earlier model trial has been rerun or regraded.

The next disability-addition increment adds two already captured whole pages and 14 support relationships. Its semantic index has 903 records, 1,464 assertions, 98 selected source pages and 29 support dependencies. It distinguishes actual carer-benefit payment from the specified disability-benefit receipt qualifications and preserves the partner-only patient scope. The care-home profiles investigate possible branches without assuming that a claimant has no partner. All 203 obligations remain open.

Its recorded 40-case development evaluation and deterministic replay pass, retaining 177 of 177 known candidate occurrences at the default 512 KiB budget. The larger care-home packages retain 61 records and 124 relationships at that budget. At 256 KiB they retain only 15 records and explicitly report missing qualification evidence. This is a measured limit, not an answer-quality score. Independent source review and 66 focused source/Reader controls pass; the new source's full publication and compact delivery remain separate work.

The partner qualification follow-up adds ten further dependencies, with conditional paths for Staff 012, 013, 014 and 017. That separate index has 903 records, 1,482 assertions and 39 support dependencies. Its separate comparison retains all 585 declared path occurrences at 512 KiB and 449 at 256 KiB, with 177/177 candidate overlap at both sizes. The care-home packages contain 55 records and 115 relationships at 512 KiB and remain insufficient. Retention does not establish household facts, legal applicability or entitlement.

The ignored-person and normal-residence increment adds two conditional concepts and eight newly selected whole pages. The current index has 913 records, 53 concepts, 1,526 assertions and 61 support dependencies. All 520 previous evidence records and 203 open obligations are unchanged. Independent source review and 61 focused controls pass. The broader declared requirements exceed the tested context budgets: candidate discovery still finds 177/177 at 512 KiB, but required-path retention is incomplete. The separate frozen comparison records this regression; it must not be presented as a complete evidence package. The Monday service/trial candidate remains fixed to the separately reviewed partner source while this larger closure is investigated.

The earlier public service 0.4.0 used the staff corpus at 9de52acf…. Its 0.4.0 SDK observation passed five exact-engine cases and compact reconstruction for all three source versions. The public staff reader check passes functional evidence checks in all three engines. Chrome and WebKit pass strict console checks; Firefox retains hosting cookie warnings (BL023). Those observations remain unchanged. Use the Monday handover for the demonstration, and the shared status page for the latest recorded release. The 0.6.0 observation preserves five approved source versions with explicit engine compatibility.

Preserved published baseline: evidence review and DMG navigation

Open the verified DMG navigation or follow the compact evidence demonstration. The public Chrome observation verified benefit, circumstance, topic and authored-concept filters across Reader, Graph and Timeline, with matching counts of 74, 343, 268 and 1. It checked all 21 application files and 188 observed corpus files, capture/source date roles, and the Timeline display limit. No console errors or targeted accessibility violations were found. This is a dated, scoped check, not whole-site conformance.

The compact-delivery guide preserves the earlier service 0.3.1 and its official SDK acceptance: a small catalogue, exact bounded reads and a replay link, while preserving the original full-package tool. The 31,312-byte abroad package reconstructs exactly and stays insufficient. That corrected public reader's twelve historical-profile journeys and a separate full-corpus Chrome journey passed their functional checks, but their strict console gates failed because CSP blocked a host-injected script; Firefox also reported cookie-domain errors. The full-corpus journey made four real tool calls and verified rendered source and diagnostic hashes on hosting version 7. The corrected service's live SDK receipt retains exact package parity. Earlier version 6 observations and failures remain intact. The recorded browser limitation and DWP-BL-023 remain open; this is not an overall public-browser pass.

Independent final review caught a stale replay link after changing or resubmitting a question. The correction passed twelve local browser journeys across three engines. The final build also exposed an Explorer documentation-cache dependency gap: linked service pages and their exact Markdown now participate in the cache identity. Neither correction changes source evidence or establishes answer quality.

The new review descriptor adds benefit, circumstance, topic and existing-concept navigation to the Reader. Its classification manifest accounts for all 19,090 captured pages, including ADM and explicit unclassified values. The 44 authored labels describe literal discovery categories; they do not establish which benefit rules apply. The Reader projection retains its DMG scope while the classification audit and Ask context cover both captured manuals. The additive combined Reader now includes ADM and cross-manual navigation, with local and public browser checks retained under DWP-BL-024. The public observations retain exact scope and limitations.

Forty staff review packs give readers small, source-linked starting points: 3,795–12,062 bytes per machine-readable pack, 42 shared evidence resources and exact bounded source excerpts. They distinguish independently located candidates from pages retained by the recorded retrieval run. All 40 staff-question occurrences and all 43 full-corpus evaluation cases remain insufficient and await evidence closure and specialist approval of the new profiles. Paired staff trials retain 49 claims and 79 citations, including quotation failures and missing qualifications; independent specialist review remains open.

A separate actual Claude client observation records seven local compact-tool calls and exact replay of diagnostics plus two source records. It also preserves the earlier no-call failure, where the model invented a catalogue while tools were disabled. Neither is an answer-quality or public-deployment claim.

Explorer PR 125 is merged and its published DMG navigation is verified. DWP PR 10 records integration of the evidence and documentation; consult that PR for its exact validation and merge status. Existing demonstrations below retain their original source and consumer versions. The work log and 24-item backlog retain incomplete work explicitly.

Current source and question coverage

The two acquired manuals contain 513 PDFs and 19,090 measured pages. Acquisition preserves documents; it does not establish complete policy modelling or specialist acceptance.

Source family Captured PDFs Measured pages Pages with nonempty extracted text Pages with no extracted text
DMG 15 September 2026 331 14,743 13,941 802
ADM 19 September 2026 182 4,347 4,256 91
Total Separate frozen observations 513 19,090 18,197 893

Every page remains represented, including the 893 pages with no extracted text. Those pages retain original PDF links; no OCR or whole-corpus visual review has established whether they are blank or image-only. Nonempty text can still contain extraction defects. Original capture dates and source publication dates remain separate.

The staff-question registry contains 40 occurrences and 39 distinct questions. The recorded remote baseline tested every occurrence against the preserved 52-record custody profile: all 40 packages matched the Explorer engine, and all 40 were insufficient, without assembly-budget truncation. Forty-two source candidates were separately verified. These are honest coverage gaps, not 40 answered questions or expert-approved benefits conclusions. No token or monetary saving has been established.

The historical 19 September full-corpus remote run tests all 40 questions and three boundary controls against the combined manuals. All 40 staff questions return candidate evidence, including ADM pages. Twelve retain an independently located candidate page and 21 retain a page from a candidate document. These measures show where retrieval helps and where it needs work; they are not answer-quality scores. All 43 results remain insufficient, with no AI answers or specialist acceptance. All 43 complete remote packages match the shared Explorer engine. Separately, official SDK calls to the deployed service match the shared engine for current imprisonment, current hospital and the explicitly selected historical imprisonment case. An actual bounded ChatGPT call inspected six source pages about international issues: 31,312 bytes, no reported host truncation, still insufficient. The published Explorer and native WebMCP check also passed: its UI, browser tools and remote service returned the same complete bounded package. This does not establish complete answerability or Voice support. Independent replay verified all 43 packages and 42 source candidates; all 21 corruption controls were rejected. The before-and-after comparison preserves the initial retrieval flaw and the improvement from 10 to 12 exact-page overlaps and 19 to 21 document overlaps; neither is an answer-quality score. The staff-question results and Monday walkthrough compare the two runs and explain the remaining work.

Full-corpus retrieval ranks literal question terms and takes bounded whole-page candidates. Existing concept aliases and relationships keep their original scopes, including custody-specific concepts. A lexical match does not establish which rules apply or that every necessary exception is represented. The immutable baseline receipts below retain their original scope. CPAG remains an external reference, with no handbook-body acquisition or reuse permission established.

The relationship audit separately distinguishes authored links preserved in the data, sparse domain modelling and Explorer display defects. The repository governance record records required pull requests, the validate merge check and private-input protections.

Ask OKF: governed context assembly

The recorded 0.6.0 service uses both captured manuals and the partner-qualified household/statutory semantic layer at content version 723bcc5b015ab38a026625c2148edbd784edf7c7. Open that fixed source in Explorer or follow the Monday handover. A fixed source link does not freeze the Explorer application; its local Reader checks and the public compact-service checks have separate receipts.

The preserved full-corpus baseline uses the additive corpus descriptor, at content version bf50ef8d91b9f1ccc2cbdb354198eae74c9ed752. Its five-minute presentation and bounded ChatGPT rehearsal remain available as historical demonstrations. Its 43-case remote evaluation, three-case live SDK verification, bounded ChatGPT observation and published-browser journeys are recorded. The general corpus returns candidate evidence and explicit insufficiency, including for imprisonment; it does not inherit a completeness claim from the older custody profile.

Historical published Explorer and browser-tool checks

Open the verified full-corpus demonstration. On 19 September, Explorer commit a8628fdb77c1c03a5d99b6d105d9e4b8722088d7 passed the recorded public journeys; downloaded application files matched the tested build. Search showed 219 imprisonment matches, with 200 displayed. The default imprisonment task resolved eight concepts and assembled 64 records, 127 relationships and 516,146 bytes. Chapter 12 routing to chapters 24, 53, 54 and 78 was visible; the package remained insufficient and truncated.

For “What happens to your benefits if you go abroad?”, set Ask's Package bytes to 32768. The six-source-page, 31,312-byte package matched native WebMCP build and explain calls and the retained remote package by complete canonical content. Its context identifier starts 2cdfa5fe; the guide records the full identifier and comparison. Native WebMCP was verified in the Codex in-app browser; ChatGPT Voice and room audio remain untested.

Preserved custody acceptance case

The following checks use their immutable earlier version. Its 52-record evidence profile is narrower than the acquired manuals, and its scoped sufficiency does not describe the new default corpus.

Open the public Ask OKF demonstration, then select Ask OKF and use the exact question in the five-minute demonstration script. It assembles a traceable evidence package for JSA, Income Support, State Pension Credit and ESA. It resolves declared concepts, traverses directed source relationships and checks whole evidence, provenance, scope and budgets before reporting scoped sufficiency. The independent A–H acceptance case checks source and graph identities rather than keyword overlap. It is separate from the earlier 160 model answer trials. No AI answer or individual entitlement decision is generated; specialist and native tool-host acceptance remain separate gates.

The public browser check verified 52 context records, 127 relationships, chapter routing, source provenance, inspectable JSON and an insufficient result under a constrained byte budget. The package matched a fresh direct engine run. DWP's 56 unit tests and all context controls passed; the unchanged earlier answer trials remain separate. This context release uses content commit efb05c66616a9cd4328a86cf412780fe7bc7cf0b; the full-corpus and pilot links below preserve their earlier release scopes.

Remote Ask OKF acceptance

The remote MCP adapter lives in OKF Explorer and imports its existing context engine. This repository supplies the remote acceptance cases and replay instructions, including the hospital question. The hospital coverage review distinguishes material acquired in the wider corpus from evidence governed by the pinned Ask custody index. Hospital evidence in that profile is insufficient; retrieved custody records must not be presented as hospital guidance. Deployment and actual ChatGPT invocation are separate gates recorded in the remote demonstration guide.

What is here

The full-DMG research candidate is publicly available at an immutable, browser-checked content commit. Try the full corpus or follow the full-DMG walkthrough. Publication and canonical CI history are tracked in PR 5; the immutable browser receipt applies to the content commit named below. The owner authorised unattended processing against the frozen 331-PDF census, with a 24 September content freeze for the 30 September seminar. The completion plan and current checkpoint record the acceptance gates. The earlier links retain their original Pension Credit scope.

The full-DMG acquisition contains 331 PDFs and 14,743 measured pages, with original bytes, source hashes and extraction-quality flags. Eight wider authoring batches add 245 concepts and 328 source-backed proposals across 78 of 78 substantive PDF units. Including the pilot, the candidate has 271 concepts and 343 semantic proposals, within 15,390 entities, 16,210 assertions and 15,363 resource records. Measured coverage distinguishes selected-passage research from complete policy modelling. Specialist acceptance remains zero.

Layer Full-DMG research candidate Preserved Pension Credit pilot
Frozen sources 331 PDFs; 14,743 measured pages Original 36 PDFs; 1,524 pages
Searchable content Every extracted page, including labelled memos, amendments, transitional and reference material Seven substantive chapters; 744 default page records; other acquired roles require explicit inclusion
Authored concepts and proposals 271 concepts; 343 proposals, including the pilot 26 concepts; 15 proposals; eight personas, ten stories and 14 questions
Semantic delivery YAML-LD control, bounded JSON-LD/RDF shards and lazy indexed records Small YAML-LD/JSON-LD bundle and Markdown records
External adviser reference CPAG metadata and links only Same access and rights boundary
Reproducibility Locked dependencies, source/output hashes, deterministic build, consumer and evidence checks Original source identities, routes, frozen domain profile and immutable demonstration links retained

All 253 remaining source-family units have a recorded research outcome: 242 completed with documented gaps and 11 spare units marked not applicable with evidence. This is bounded accounting, not full-body review or a finding that memos and annexes contain no substantive rules. See the source-family review.

The 160 observed, context-aware, source-guided answer trials have separate model assessments: 90 supported, 56 partial and 14 with underspecified rubrics. These are assessor categories, not an accuracy score or a blind end-to-end retrieval benchmark. Original prompts, reads, responses, omissions and assessments are retained in the trial summary. Indexed retrieval controls are separate evidence; no specialist approval, comparative model result, voice or WebMCP performance is inferred.

The capture date is 15 September 2026. The original Pension Credit landing page reported 20 July 2026 as its latest update. Neither capture nor listing dates establish a provision’s current applicability. The original pilot coverage and original extraction report retain their 36-PDF scope; full-DMG coverage describes the larger candidate.

Try it

Open the full-DMG candidate at content commit 80b6f08426aea39dd2934fb8795b61215e2cc0ad, snapshot dwp-full-dmg-2026-09-15-2dd78242297c:

The public browser receipt records the exact content and consumer identities. Checked journeys include both entry formats, source text and a direct PDF link, labelled Graph relationships, PDF resource cards, CPAG's month-precision Timeline and bounded search controls. The observed ESA Graph had 11 nodes and 10 edges with no missing labels; 84351 returned six labelled PDF resources. These checks establish those interactions, not every route or policy interpretation. See the ten-minute full-DMG walkthrough.

Explorer can expand an unmatched search string into an indexed term: unavailableclaimantdetails returned seven results for unavailable in the public check. Inspect the displayed query interpretation and the source evidence. The nonsense control zzzxqvnomatch returned an unmatched term and no results; the exact Python retrieval controls test a separate path.

Launch the original meeting demonstration. It uses the immutable captured bundle and starts with the capital-disregard question and matching evidence. Browser verification records the exact snapshot and journeys. You can also load bundle/okf-bundle.json as a local file in Explorer.

Start with Where does the guide explain capital disregards? (question/pc001). Follow the question to chapter 84, PDF page 35, inspect the original source and its hash, then explore the relationships. Use question/pc003 to examine historical context and extraction limits, and question/pc008 for the scope boundary.

An AI can read the repository or the generated bundle directly. The AI guide provides a prompt requiring chapter/page citations, source hashes and explicit missing-evidence handling. Loading the files does not create an MCP server or connect to DWP systems.

Build and verify

Install uv and run:

uv sync --locked
uv run --locked python scripts/check_semantic_contract.py
uv run --locked python scripts/build_bundle.py
uv run --locked python scripts/build_bundle.py --check
uv run --locked python scripts/validate_bundle.py
uv run --locked python domain-profile/validate.py
uv run --locked python scripts/evaluate_queries.py

For the full-DMG candidate:

uv run --locked python scripts/build_full_dmg.py
uv run --locked python scripts/build_full_dmg.py --check
uv run --locked python scripts/validate_full_dmg.py
node --experimental-strip-types scripts/check_full_dmg_consumer.mjs
uv run --locked python scripts/report_full_dmg_coverage.py --check
uv run --locked python scripts/evaluate_full_dmg.py --check
uv run --locked python scripts/reconcile_full_dmg_references.py --check
uv run --locked python scripts/verify_full_dmg_trials.py --check

Its Explorer entry point is full-dmg/okf-explorer.json (or full-dmg/okf-explorer.yamlld). full-dmg/okf-bundle.yamlld is a semantic control document pointing to bounded graph shards; it is not a small-bundle import. The launch links above name the immutable candidate checked in the browser. The original 15 September release recorded 43 passing local tests. The trial verifier checks retained evidence and reproduces the assessment summary; it does not run new model answers or assign grades.

For the separately acquired ADM manual, staff-question baseline and publication guards:

uv run --locked python scripts/acquire_adm.py --check
uv run --locked python scripts/test_acquire_adm.py
uv run --locked python scripts/build_context_discovery.py --check
uv run --locked python scripts/build_context_corpus.py --check
uv run --locked python scripts/audit_relationships.py --check
uv run --locked python scripts/check_private_inputs.py
uv run --locked python scripts/test_private_inputs.py
uv run --locked python scripts/check_browser_evidence.py
node --experimental-strip-types scripts/evaluate_staff_questions.mjs --check --explorer-root ../okf-explorer
node --experimental-strip-types scripts/test_staff_questions.mjs --explorer-root ../okf-explorer

The staff-question replay needs the matching Explorer context implementation; the retained receipt binds its exact file hash. It replays saved remote responses without calling a model or a live endpoint. The private-input check rejects tracked or staged .email.md files without reading their contents; run it before committing or pushing, not only in CI. The governance guide explains these boundaries.

The browser-evidence check validates the retained public observation against the deployed application identity and the bounded MCP package. It runs offline: it does not open a browser, repeat the journeys or certify a later deployment. Actual browser journeys remain a separate required publication check.

The consumer check uses Node 26.7.0 in CI and unmodified, hash-pinned Explorer validation code under profiles/explorer-runtime/. It verifies the research notice and every route label against the actual consumer contract. Publisher titles retain their original bytes; display labels normalise whitespace only.

The reference register accounts for literal DMG paragraph, memo and legal-reference candidates across all 331 sources. It distinguishes single candidate locations, ambiguous matches and unresolved identifiers. A match supplies navigation; it does not establish legal identity, incorporation or present applicability.

These commands use the frozen source snapshot. They do not redownload the collection. Build --check rebuilds in memory and compares generated bytes with retained artefacts; the other checks validate their named evidence and projections.

Verify the additive source-led candidate

After installing the locked dependencies, these checks replay the retained PDF structure observations and compare generated outputs. They do not invoke Poppler, download sources or regenerate the earlier bundle and corpus releases:

uv run --locked python scripts/build_pdf_structure.py --check
uv run --locked python scripts/build_manual_guide.py --check
uv run --locked python scripts/build_structured_units.py --check
uv run --locked python scripts/build_structured_context.py --check
uv run --locked python -m unittest discover -s scripts -p 'test_manual*.py'
uv run --locked python -m unittest discover -s scripts -p 'test_pdf_structure*.py'
uv run --locked python -m unittest discover -s scripts -p 'test_structured*.py'
node --test scripts/test_structured_context_metrics.mjs
node --test scripts/test_report_structured_context.mjs

The source-led guide gives the separately versioned context-comparison procedure and its exact Explorer consumer. A discovery card is a small navigation record pointing to the underlying source unit; its preview is not evidence. The PDF observation notes explain why original trees, failed observations and bounded parser replays are all retained. Fresh PDF observation is an explicit maintenance operation, outside normal CI. Missing structure or an unresolved boundary remains visible.

Continuous integration (CI) means the automated checks run for each proposed change. The existing replay and the new source-led checks run in parallel jobs, each with its own resources. The protected validate check succeeds only when both jobs succeed. The learning website still publishes only after the complete workflow passes for the exact main-branch commit.

For keyword retrieval with JSON citations:

uv run --locked python scripts/query.py '84351' --limit 5
uv run --locked python scripts/query.py 'part-week payments' --limit 5

--include-history explicitly includes all acquired document roles within the selected inventory. --scope full-dmg selects the 331-PDF inventory; the default remains the original Pension Credit inventory. Results identify the exact inventory hash and source role. Broad terms can rank scenario-specific examples above general rules; the CLI is retrieval, not legal reasoning. No-result queries return no invented evidence.

A deliberate source refresh is separate: inspect scripts/acquire_sources.py --help and source instructions, acquire a new snapshot, review changes and repeat discovery and assurance. Do not silently overwrite the meaning of an existing release.

Authoring and format boundaries

OKF 0.2 is the Markdown core. “OKF+” here means that core plus the additive Explorer Bundle Wiki semantic profile; it is not a separate universal OKF standard. YAML-LD 1.0 remains a W3C Working Draft. The build uses pinned local contexts and URDNA2015 RDF normalisation; it does not claim RDFC-1.0 or SHACL conformance.

What the exemplar does not establish

The source PDFs contain historical examples, dates, scenario-specific treatments and references to memos, legislation and case law. These have not been exhaustively consolidated or legally reviewed. Machine extraction has known defects, particularly letter spacing in chapter 83. Across the 331-PDF capture, 802 pages have no extracted text; no exhaustive visual review or OCR has established whether those pages are blank or image-only. The original 36-PDF pilot had 85 such pages, including ten within its default 744 page records.

Source acquisition and bounded research attempts are complete for the declared full-DMG scope. The public browser receipt covers its named interactions. The full Foundry production gate sequence, comprehensive accessibility assurance, expert legal review and automatic legal rule modelling are not complete. Repository status records the candidate boundary and validation evidence. No claimant case data was acquired.

Reuse and next steps

Open the CPAG reference in Explorer.

The beginner research guide and file-level index separate the public CPAG contents navigation, unvalidated UC evaluation starter, client-access conversations and historical calculation design prompt from the authoritative source and published bundle.

The CPAG handbook access review explains the new searchable external reference and the permission required before handbook content could be processed or redistributed. Search CPAG in the updated bundle to inspect it. The original meeting link above preserves its earlier snapshot; the CPAG reference requires the updated bundle.

Contains public sector information licensed under the Open Government Licence v3.0. Original project code and original material use the MIT licence; source extracts retain Crown copyright, attribution and applicable exceptions. See NOTICE.md.

The later-source log records tribunal decisions, benefits calculators and CASA for later assessment. They are not incorporated into this snapshot. The product backlog records wider persona/journey evaluation, a benefits engine and application/change-of-circumstances journeys. The recorded source-guided trials do not deliver those operational journeys.

Semantic stage two and seminar preparation

Read the next-stage assessment and delivery plan. The preserved Pension Credit pilot has 26 individual concept files, 15 source-backed semantic proposals, a non-executable capital-disregard review candidate, eight personas, ten stories and 14 questions. Its semantic map retains that bounded scope. The full-DMG candidate adds the wider concepts and observed trials described above. Model assessment does not make either layer specialist-reviewed.

All 331 DMG PDFs are acquired and indexed. The separate 182 ADM PDFs have now also been acquired with original bytes and page text, giving 513 documents and 19,090 measured pages across the two source families. The 43-case combined-corpus remote evaluation and three live SDK cases pass; ChatGPT inspected the bounded abroad package. Published Explorer and native WebMCP journeys also passed for the recorded version and application commit. The earlier custody profile and its receipts have not been rewritten. The 24 September plan separates source acquisition, semantic coverage and expert review. Content freezes on 24 September for the 30 September seminar; website, WebMCP and Mac voice/audio feasibility are tracked separately.

Concept authoring now uses knowledge/**/*.yamlld. Additional checks:

uv run --locked python scripts/evaluate_semantics.py
python3 source/discovery-2026-09-15/acquire_metadata.py --check

The earlier Explorer links remain pinned to their original snapshots. Stage-two browser verification and publication state are recorded in repository status.

Open the stage-two semantic graph in Explorer. The pilot browser receipt binds this graph and the corrected CPAG Timeline to the checked content commit; the earlier semantic receipt records the detailed relationship journeys.

The additive corpus reading-help catalogue accounts for all adopted DMG/ADM structured passages and exposes bounded, hash-bound source segments and candidate annotations. Run PYTHONDONTWRITEBYTECODE=1 python3 scripts/build_reading_help_corpus.py --check to rebuild and verify it from frozen inputs. The earlier 12+24 first-page controls are preliminary integrity smoke checks. A separate source-led initial gate passed 12/12 after its first failure was retained. The first fresh held-out gate remains failed overall, with 23 passes and one failure, because its H05 expected label omitted a printed range qualifier; a revised, separately reviewed fresh 24-case source gate passed 24/24 across both manuals, abbreviation scope, continuing tables, exceptions and unresolved references. These are deterministic source controls, not specialist acceptance or a staff-answer quality gate. This remains unreviewed reading assistance, with extraction gaps and uncertain references visible.

The reading-help release record binds the implementation, checks, public verification state and retained gaps. Use the five-minute walkthrough for the presentation.

The bounded Chapter 60 reading aid gives occurrence-specific help for paragraphs 60025 and 60033. Its generated manifest keeps the frozen source, abbreviation expansion, local citation and unreviewed explanation separate. It does not alter the combined Reader programme or establish legal answerability.

Demonstration learning paths

The demonstration learning programme adds 12 paths and 112 activities to the combined Reader, with assessed prerequisites and inspectable evidence. It is independent teaching material, not official DWP training.

The bounded reading-help legislation bridge reuses the existing statutory evidence and retains three further Chapter 60 targets with separate geographical versions. The rollout work log records engineering and publication gates; candidate overlays do not silently change deployed evidence.