Research Payload Family
Research Payload Family shared core v0.1.0
Payload candidate for joining pre-release evaluation, field testing, deployed monitoring, change detection, recourse, and continue–revise–pause–rollback decisions across a named context of use.
v0.1.1
Candidate state: this page defines the research job, parameters, evidence spine, method skeleton, outputs, and promotion test for a reusable evaluation and monitoring payload. A completed cross-layer specimen and run receipt remain the evidence needed for executable status.
Research job
Support a named decision about whether an AI-enabled system should continue, expand, change, narrow, pause, roll back, or retire within a declared context of use.
The candidate joins four evidence horizons that are often separated:
- pre-release evaluation;
- field testing under bounded operating conditions;
- post-deployment monitoring;
- reassessment after material change, incident, or population shift.
The result should show which claim each observation supports, where detection remains limited, and which accountable receiver can act.
Why this candidate is timely
NIST AI 800-4 — Challenges to the monitoring of deployed AI systems, published 2026-03-06, distinguishes controlled evaluation from the need to observe systems in real deployment contexts. The report describes monitoring categories while finding that methods, validated practice, and shared terminology are still developing.
That evidence supports a configurable local research instrument. The NIST report orients monitoring design; the named receiving context supplies its population, metrics, thresholds, recourse, and operating authority.
Required parameters
Parameter | Required entry |
System and version | Model, software, prompt, tool, data, workflow, and deployment version |
Context of use | Task, affected population, environment, channel, and consequence |
Receiving decision | Continue, expand, revise, restrict, pause, rollback, retire, or study next |
Accountable receiver | Role authorized to act on the evidence |
Claim inventory | Capability, reliability, safety, quality, access, privacy, or operational claims |
Baseline and comparator | Current practice, prior version, control, threshold, or service objective |
Observation horizon | Pre-release, field test, deployed interval, and review date |
Data boundary | Public, synthetic, authorized, restricted, de-identified, or aggregated |
Change boundary | Events that create a new evaluation state |
Recourse | Reporting, correction, appeal, human review, and recovery route |
A missing value that changes the governing population, metric, authority, or threshold produces a scoping pause.
Evidence classes
- Controlled evaluation: repeatable tasks, fixtures, simulations, or benchmarks under declared conditions.
- Field evidence: bounded use in the receiving environment with named participants, time window, and safeguards.
- Operational evidence: production traces, service measures, incidents, corrections, overrides, and recovery.
- Affected-population evidence: accessibility, burden, access, outcomes, recourse, and subgroup experience.
- Source and policy evidence: current standards, regulation, documentation, contracts, and organizational policy.
- Interpretive judgment: source-linked reasoning whose status remains visible.
- Open local hypothesis: a claim with a defined observation, threshold, and disconfirming result.
Several dashboards repeating one underlying signal remain one evidence lineage.
Evaluation and monitoring map
Horizon | Primary question | Example observation | Decision effect |
Pre-release | Can the system perform the named task within the initial boundary? | Task accuracy, failure modes, calibration, accessibility, security, robustness | Admit to field test, revise, or stop |
Field test | How does the system interact with real workflow, people, and environment under bounded exposure? | Workload, corrections, handoffs, comprehension, recourse, subgroup signals | Expand, narrow, revise, or pause |
Deployed monitoring | Does operation remain within the accepted envelope? | Drift, incidents, overrides, outcome and service measures, data and tool change | Continue, investigate, restrict, or roll back |
Material change | Which prior evidence still travels after a change? | Version diff, dependency change, population shift, new task, new regulation | Reuse bounded evidence or require reassessment |
Recovery and retirement | Can the system return to a coherent prior or alternative state? | Rollback, data reconciliation, preserved access, handoff receipt | Recover, retire, or transition |
Research demands
1. Context and consequence
Define intended use, foreseeable adjacent use, affected populations, material consequences, and the role that retains decision authority.
2. Claim-to-evidence map
For every material claim, record the fitting observation, source class, population, environment, time horizon, acceptance observation or decision rule, and coverage boundary. Every result also declares its observation standing as observed_real, observed_synthetic, planned, or unavailable; every threshold declares its provenance as source_prescribed, organization_selected, or exploratory.
3. Change inventory
Identify changes to model, prompt, retrieval corpus, tool, software, policy, data, population, interface, workflow, and dependency. Classify which changes preserve evidence and which reopen evaluation.
4. Detection design
For every monitored signal, record denominator, observation window, collection method, detection limit, alert threshold, false-positive and false-negative consequence, owner, and response. Review coverage across the NIST AI 800-4 categories of functionality, operational conditions, human factors, security, compliance, and large-scale impacts; an open category remains visible with its fitting next observation.
5. Human review and recourse
Observe review workload, override and correction behavior, escalation, explanation, appeal, access, and the path back into source or system change.
6. Counterevidence and benign explanations
Search for alternate causes such as workflow change, data latency, coding change, seasonality, instrumentation failure, selection effects, or reviewer adaptation before fixing a causal interpretation.
7. Continue–change–pause criteria
Define evidence thresholds and the authority for continue, expand, revise, restrict, pause, rollback, retire, and study-next decisions.
8. Longitudinal receipt
Preserve system identity, evaluation version, operating conditions, evidence date, decisions, open gaps, next review, and events that force reassessment.
Minimum output contract
- context-of-use declaration;
- system and change inventory;
- claim-to-evidence matrix;
- pre-release evaluation plan and results;
- field-test plan and results where authorized;
- deployed monitoring specification;
- affected-population and recourse evidence;
- incident, correction, and recovery route;
- thresholds and accountable decisions;
- coverage ledger;
- completion judgment;
- run receipt and next review.
Companion instruments
- Research Payload Family — Shared Core
- Evidence-Bearing Verification Annex
- Decision & Data Shape Parameter Pack
- Goal Fidelity Under Partial Context — Deep Research Result: Evidence-Bearing Collaboration Across Model, Harness, Human & Organization
- Material Verification — Connect Checks to Behavior, Fault & Observable Evidence
Promotion work
Executable status requires:
- one worked specimen that spans controlled evaluation, bounded field evidence, deployed monitoring design, and a change decision;
- one partial or unavailable evidence case;
- one incident or recovery path;
- one run receipt;
- a completed coverage ledger;
- a review by a fitting system, domain, and affected-population vantage;
- a promotion decision recorded in the portfolio.
Until those observations exist, this page remains a public-grade candidate and a usable planning surface.
Focused refinement from dry run
The synthetic dry-run specimen and its separate run receipt exercised the candidate across controlled evaluation, field-like rehearsal, operational replay, material change, containment, repair, and recovery.
The run established control-flow coherence and exposed four focused improvements now carried by v0.1.1:
- every result declares
observed_real,observed_synthetic,planned, orunavailable; - every threshold declares
source_prescribed,organization_selected, orexploratory; - the six NIST AI 800-4 monitoring categories form a required coverage ledger alongside lifecycle horizons;
- evidence from synthetic execution and evidence for promotion remain distinct, so real field, deployed-system, affected-population, and fitting independent-review observations retain their own standing.
The candidate remains Blooming. The next promotion observation is an authorized real-world specimen with system, domain, and affected-population review.