Source-led results and remaining work
This is an independent research candidate, not official DWP guidance. The figures below describe preserved source, navigation and evidence delivery. They do not measure legal correctness or AI-answer accuracy.
What is now processed
All 513 captured DMG and ADM PDFs, 19,090 pages and 35,143,443 original extracted text bytes are accounted for. The source projection contains 53,727 units, including 75 earlier author-declared units and 10,406 units crossing page boundaries. PDF pages remain exact checking locations.
The manual guide records 21 source-supported, scoped reading conventions. Compact discovery cards point to full passages. The context graph contains 19,193 relationships, including 17,952 resolved literal paragraph references. It also records 30,057 unresolved reference observations. An unresolved reference is retained evidence about the build's limits, not a guessed link.
The full-corpus check verifies every card/unit pair, all 67,087 source spans, the complete relationship index and all 1,355 input bindings. It does not review the legal interpretation of every passage. Machine-detected boundaries remain explicitly unresolved for completeness.
What the staff-question experiment shows
Trial 06 keeps the 40 original question occurrences, including duplicates, plus an unknown-question control. Two projections and two byte budgets produce 164 assemblies. It makes no network or model calls. The engine, questions and ranking parameters are fixed; this increment changes source selections.
| Measure | Earlier logical-unit baseline | New structured corpus at 512 KiB |
|---|---|---|
| Staff occurrences activating a source-read profile | 8/40 | 37/40 |
| Staff occurrences retaining source evidence | 40/40 | 40/40 |
| Staff occurrences retaining relationships | 18/40 | 40/40 |
| Retained inherited source-selection paths | 47/47 | 47/47 |
| Retained previously added source-selection paths | Not declared | 181/181 |
| Retained newly added source-selection paths | Not declared | 208/208 |
| Results marked insufficient | 40/40 | 40/40 |
Path totals are case incidences: one shared route may count in several questions. A source-read profile is an explicit proposal about passages worth checking. It is not a complete legal answer rubric. All three path groups must stay separate when comparing earlier trials.
At 32 KiB, the new corpus retains no source evidence in any of the 40 staff packages. This is a deliberately smaller-budget control; the current Explorer's assembly default is 512 KiB. Explanations and unresolved obligations consume the smaller budget. A smaller discovery card must not stand in for omitted evidence. Use a larger assembly with bounded exact delivery when the compatible service release is admitted. That separates total selected evidence from the size of each reply.
The three remaining source-read profile gaps are:
- Staff 001: benefits abroad, without a named benefit or absence facts.
- Staff 008: ambiguous SDA terminology and relevant dates.
- Staff 018: an unnamed benefit and multiple care-home possibilities.
Existing leads remain available. The system must request or expose those distinctions before treating a particular benefit route as applicable.
Preserve the less favourable measures
The 40 original page-based research requirements and 203 open obligations are unchanged. Location overlap with a new unit does not satisfy an old page requirement. Only 52/776 original page-path incidences are retained.
The new allocation retains 469/765 restored location-route incidences and 34/87 inferred route incidences. Trial 05 retained 490/765 and 35/87 respectively. These navigation trade-offs are explicit; the complete retention of source-read paths must not hide them. All large staff results report both retrieval and assembly truncation.
The original imprisonment, hospital and unknown-question cases have a separate evaluation. They are not part of the 40-question denominator. Hospital remains an insufficiency control. Restored custody navigation preserves the earlier requirements and unresolved paragraph selectors; it cannot silently establish modern benefit applicability from a historical manual.
The first separate imprisonment run exposes a delivery failure: at 512 KiB it retains 20/53 selected source units and 43/89 declared paths. The text of 12002, 12003, 12015 and 12016 is present, but none of the four complete Chapter 12 routing paths survives and the supporting relationship from 12003 to 12002 is omitted. This is being investigated separately. The 40 staff results must not be used to imply that this original demonstration case passes.
What remains
- Independent review of boundaries and legal applicability, including 28 large paragraph candidates retained for source triage.
- Resolve the three underspecified task families without guessing facts.
- Review and close individual evidence obligations against authoritative, appropriately dated source material. None is closed by this census.
- Measure answers against the same fixed evidence, with claim-level support and qualifications. Earlier model trials remain historical; this increment does not rerun them or establish better answers or affordability.
- Complete exact publication and client observations. A merged engine, published corpus and observed external-client call are distinct milestones.
Check the records
- Trial 06 machine-readable report
- Generated comparison
- Full-corpus integrity receipt
- Original acceptance restoration
- Beginner explanation
- Demonstration script
- Backlog and named acceptance checks
These are known development questions. No blind evaluation, specialist acceptance, individual award decision or complete legal coverage is claimed.