OKF-DWP

Independent experimental publication. Not an official DWP service, benefits advice or an entitlement calculator.

View this version’s Markdown source

On this page

Use the Data agent to propose a semantic review

Learning path · Portable methodology · Check the ChatGPT connection

A first frozen-file review has returned proposals. This guide records its limited observations and the repeatable review method. It does not establish an automatic ontology feature or a completed accuracy benchmark. OKF-DWP is an independent experiment, not an official DWP service or benefits advice.

What the documented workflow adds

A connector gives an AI access to permitted information and operations in another system. An analysis workflow helps it decide what to investigate, check the inputs and produce a useful result. OpenAI's documentation separates these server-backed tools from skills: reusable instructions for doing a task. Plugin architecture

The documented Data Analytics workflow in ChatGPT Work covers gathering source context, checking missing values and duplicates, joining datasets, exploring hypotheses, modelling, validating and producing reports, notebooks, charts or dashboards. It asks for visible assumptions and limitations, with original inputs preserved. This is more than connecting to a database. Analyse datasets and ship reports

The product announcement of 10 September 2026 explicitly names the Data agent in ChatGPT Work, listed as Data in the Plugins directory. It can investigate questions, use organisational definitions and data relationships, combine permitted data and documents, explain evidence, and produce or refine interactive dashboards. Connected actions require their own authorisation. It consumes trusted semantic context; the announcement does not promise automatic legal ontology generation. Data agent announcement

The January 2026 article about OpenAI's in-house data agent describes a separate custom internal system. Its layered context, human annotations, code-derived meanings and evaluation are useful design lessons, not proof that every internal capability is exposed in the public Data plugin. Internal engineering account

An ontology defines a domain's kinds of things and relationships. OpenAI's knowledge-graph cookbook demonstrates a custom pipeline that turns documents into statements, dated relationships and resolved entities. That supports the feasibility of model-assisted proposals. It is an engineering example with June 2025 benchmarks, not a promise that ChatGPT Data Agent supplies that pipeline or understands a whole corpus. Temporal knowledge-graph cookbook

First parallel review, 21 September 2026

The separate ChatGPT Work task Semantic quality review returned a report and JSON proposal artefact for frozen source commit 723bcc5b015ab38a026625c2148edbd784edf7c7. Its returned account says it used the Data Analytics report workflow, local Git, structured-file inspection, hashing, PDF rendering and visual inspection. Direct raw-file web reads failed with DisabledError; it obtained the exact public commit through Git instead.

It reported no callable Ask OKF tools and no separate nested Data Agent call. This is a frozen-file semantic review, not a live MCP acceptance test. The local implementation task inspected the returned messages and proposal summary; it has not imported or independently verified the two complete cloud artefacts. The following file identities are therefore reported checksums, not local verification receipts:

Reported artefact Reported SHA-256
OKF-DWP_semantic_quality_review_723bcc5_2026-09-21.md 0da1f5e61bb6ec0088a194db087e9c151d25695cad3287d07e3e82bd8ce280db
OKF-DWP_semantic_proposals_723bcc5_2026-09-21.json baa71e73d698736a4f174bf378695b9956af67626e78a30764e22b4435a385cb

The useful contribution was comparison across the source inventory, extracted pages, concepts, task profiles and catalogue. It classified defects and proposed source-linked relationships, review gates and tests. No database connector was needed for that exercise.

Proposal awaiting review Source locator in the frozen extraction Why it matters
Qualifying young person DMG 077008–077014, Chapter 7 Part 6, PDF pages 12–13 Keep household conditions and further qualifications together.
Expected absence duration DMG 077001 and its examples, pages 8–9, with related branches on pages 9–12 An expected duration and elapsed time are not interchangeable facts.
Absence purpose DMG 077003 and 077007, pages 9–12 A general travel mention cannot establish the purpose required by a passage.
Passage continuation dependencies Pages 8–9, 9–10, 10–11, 11–12 and 13–14 A page boundary must not cut an example or qualification away from its support.

The review also recommended reusing the existing international concept IDs and keeping the general abroad question as a clarification and routing task. Those points agree with the independently developed bounded repair. Agreement between model-assisted reviews is not specialist acceptance. The additional proposals remain model-derived and unreviewed; they have not been inserted automatically into the semantic source. Current legal applicability, DMG 04642, cited regulations and relevant ADM evidence remain unresolved.

Before integration, obtain the full artefacts, verify their bytes and exact quotations, review every condition and date limit, and test the proposed change against fixed inputs. No held-out benchmark, improved AI-answer accuracy, affordability or complete benefits coverage has been established.

A small proposal-and-review trial

Use the existing discovery-first method, rather than starting another vocabulary from scratch. A concept is a named meaning; an assertion links records in a stated direction. Provenance records where that assertion and its supporting material came from.

  1. Fix the question and inputs. Select one topic and an initial 10–20 public source pages, including headings and necessary continuations. Record the source versions, file hashes, page identifiers, rights and exclusions. A hash identifies exact bytes. Log missing supporting pages; do not fill their content from memory or quietly expand the trial.
  2. Supply the existing model. Give the agent the relevant concept catalogue, vocabulary guide and task requirements. Reuse stable identifiers where meanings match. Keep benefit variants, components, entitlement and payment distinct. Similar names do not establish identical meanings; an ambiguous abbreviation must retain its alternatives.
  3. Request proposals with evidence. For each proposed concept or directed relationship, record its definition or meaning, existing or proposed identifier, source quotation and locator, supporting heading, conditions, exceptions, territory and date limits. Keep source publication, capture and applicability dates separate. Use existing predicates where their meaning fits: a reference is not legal applicability; dcterms:requires identifies supporting material that must accompany an interpretation.
  4. Keep an explicit gap list. Mark each interpretation model-derived and unreviewed, separately from source authority. Retain uncertainty, conflicting passages and missing dependencies. An AI's confidence is not a specialist's approval. Do not change an open obligation to complete because another model agrees.
  5. Review before integration. Return a proposal table, supporting evidence inventory, unresolved questions and the exact run settings. Accepted changes go through the established additive authoring and producer workflow in staff-question semantics. Do not hand-edit generated bundles or rewrite frozen source releases.

For example, a page mentioning a benefit can justify a candidate for navigation. It does not establish that its rule applies to everyone receiving that benefit. A useful review finds the limiting heading and exceptions as well as the term.

Three different checks

Check What it can establish What it cannot establish alone
Mechanical validation Identifiers resolve, quotations match the retained extraction, hashes agree and required fields are present That extraction is faithful to the original or an interpretation is correct
Source review The original passage, heading, continuation and qualifications support the proposed meaning within its stated scope Complete legal coverage or applicability to an individual
Specialist review A named reviewer accepts a stated interpretation and evidence requirement within a defined scope Automatic approval of the rest of the corpus

A second AI can help identify defects but is not a benefits or legal specialist. The proposal and each review decision need separate authorship and status.

Measure improvement without teaching to the test

The supplied staff questions are development cases: they helped shape this bundle. Paraphrasing them does not make an independent accuracy benchmark. Before a new trial, reserve additional questions and expert-reviewed expected evidence that the proposal process will not see. These are held-out questions.

Compare the existing and proposed semantic models with the same source pages, engine and budgets. Report counts and denominators for source-supported assertions, unsupported relationships, missed conditions or exceptions, and required evidence paths retained. Include an ambiguous term, an unavailable source, an unknown term and a budget too small to hold required support. Expected behaviour may be a clear refusal or an unresolved alternative.

Evaluate AI answers separately against the exact supplied evidence: each claim needs support and its qualifications. More graph edges, valid citations or more retrieved pages do not by themselves mean more accurate answers. Retain failed attempts and unchanged baseline results. The first proposal review has no held-out accuracy scores.

Establish access at the point of use

OpenAI directs authors to test the installed plugin's actual tool selection, arguments, results and errors; capabilities depend on the chat's tools and permissions. A plugin visible in a parent conversation does not prove that a separate Data Agent can call it. Plugin permissions · Connection and testing guidance

For a live Ask OKF trial, follow the client check and preserve source, engine, context identity, budgets and actual evidence reads. Stop on missing tools or a rejected call. Do not relabel website browsing as a successful tool invocation.

A separately supplied retained evidence package can support an offline proposal exercise. Label it a frozen evidence handoff, verify its contents where possible and retain its gaps. It is not a fresh MCP call, proof of Data Agent access or a complete representation of the source corpus. Keep claimant information and private account material out of this public trial.