# D0 → D1 → D2: proposed independent trial protocol

Status: **PROPOSED_NOT_EXECUTED**. This document is a preregistration checklist,
not a runnable package, a replication receipt or a new scientific result.
The [historical internal case](successor-case.html) is a separate record.
Its original hidden cases are not supplied, and their reported scores are
not acceptance criteria or expected answers for the new trial.

An owner-authorized first run may be an **AI-assisted technical evaluation**
by a declared AI tool. Record model/tool version, operator,
candidate authorship, access to evaluation data and all interventions. Do
not describe that review as a human expert or independent external
validation. If the same assistant context can read candidate construction
and hidden answers, blind independent judging has not been established.
An independently administered external replication is a separate milestone.

## 1. Prerequisites and authority boundary

Before beginning, the evaluator must receive an explicitly approved,
sanitized frozen package containing the component needed for this test,
permitted initial state, D1 artifact bytes, loading/ablation interface,
execution instructions, dependency lock and licenses. Record SHA-256 for
every file and a manifest binding the complete package. No live private
Resident, production memory, credential or original holdout is required or
authorized. These executable materials are **not currently downloadable from
this site**; producing or publishing them requires separate review and
authorization. Stop if the required package or authority is missing.

The evaluator controls an isolated execution environment, independent scoring
code and concealed expected outputs. The candidate cannot read the Judge,
hidden answers or scoring storage, change the rubric, use extra tools or
reach unapproved network services. Record the sandbox controls actually
enforced; do not infer isolation from a status label.

## 2. Preregister before revealing tasks

Fill and sign/date these fields; blank fields are not a completed protocol:

- Trial ID, evaluator, roles and intervention policy.
- Exact baseline/D1 package and state manifests; permitted learning inputs;
  D2 artifact schema and provenance obligations; tool/environment versions.
- New task-family definition, task-generation procedure and random seed
  custody. Commit the protocol before the candidate sees any scoring cases.
- Primary metric, minimum improvement, failure threshold, repetitions,
  uncertainty reporting and any multiplicity adjustment.
- Wall time, CPU/GPU time, peak RAM, actions, tokens, observations and network
  allowance per arm, plus how prior D1 acquisition cost is accounted for.
- Fixed stopping rule, restart definition, transfer distance and permitted
  adaptation budget. Equal execution budgets do not imply equal lifetime
  training cost; disclose both instead of hiding the D1 acquisition cost.
- PASS, FAIL and INVALID definitions, disqualifying leakage, missing-evidence
  rules and who can authorize emergency termination.

No numeric threshold or resource value is implied by this template. If the
evaluator changes any scientific criterion after observing results, freeze a
new protocol/version and fresh concealed cases; do not rewrite the old result.

## 3. Fresh evaluator-owned tasks and holdout

Only after the exact package and protocol are frozen, the evaluator creates
new tasks that were not supplied by the candidate or its engineer. Separate
learning/development observations from a concealed evaluation holdout and a
separate transfer family. Define family separation, deduplication and leakage
checks in advance; changing only names or seeds is not necessarily transfer.

Commit case IDs, partitions and expected-output/scoring hashes in evaluator
custody. Candidate-visible inputs must not reveal expected answers, repairs,
Judge feedback or hidden case structure. Keep evaluation immutable while
comparing arms. Do not publish hidden answers merely to make a download.

## 4. Execute and preserve the causal comparisons

Use independent clean state for each arm, identical task partitions and the
preregistered observation and execution budgets. Counterbalance run order
when appropriate and preserve raw outputs, failures, seeds and resource use.

1. **D0 acquisition baseline:** run the frozen component without D1 on the
   new learning observations. Record whether it authors an executable second
   artifact; do not replace failure with an engineer-supplied solution.
2. **D1-assisted acquisition:** enable only the frozen D1 artifact, present
   the same permitted observations and budget, and require candidate-authored
   D2. Bind D2 bytes and parent reference to immutable hashes before scoring.
   Distinguish candidate authorship from generic engineered orchestration.
3. **Held-out comparison:** score baseline and acquired candidate on the
   concealed cases without revealing the answers or adapting on that holdout.
   An executable artifact alone is not a measured improvement.
4. **Parent ablation:** remove only D1 using the preregistered interface while
   preserving other conditions. Measure the effect on the declared D2 claim.
   A broken loader is not automatically scientific evidence: record the
   failure mode and whether dependency loss was the preregistered endpoint.
5. **Exact parent restore:** restore the same D1 bytes and verify the hash;
   repeat the evaluation without providing extra learning evidence.
6. **Child ablation and restore:** remove only D2, measure loss of its claimed
   gain, then restore its exact bytes and repeat. Keep any failures.
7. **Fresh-process restart:** terminate the test process, retain only the
   explicitly allowed serialized state, launch a new process and re-run the
   retention checks. Preserve process-boundary and state-commitment receipts;
   a log message saying restart is not sufficient evidence.
8. **Transfer:** evaluate the separately concealed new family under the
   predefined adaptation budget, including baseline and relevant ablations.
   Repeat the parent control after restart if required by the frozen claim.

Every arm must be traceable to its inputs, artifact identities and budget.
Do not spend a larger search budget on the preferred arm, omit unsuccessful
seeds, silently change artifacts or use hidden answers to author a repair.

## 5. Independent decision and inspectable record

The independent evaluator/Judge applies the frozen criteria. Report effects,
uncertainty, resources, exact failures and every human intervention. A
signature establishes provenance under a trusted key, not scientific truth.

The agreed sanitized evidence bundle should include the preregistration,
public file manifest, artifact commitments, task-partition commitments,
per-arm results, attribution, ablation/restore and restart receipts, resource
ledger, evaluator decision and a verifier for that exact schema. Raw private
inputs or hidden answers stay in evaluator custody unless separately approved
for release. Explicitly list every omitted item and the resulting limitation.

A missing package, leaked holdout, unsupported sandbox guarantee or missing
critical receipt makes reproduction incomplete or invalid under the frozen
rules; it is not a pass. A correctly verified but unsuccessful experiment is
still a failed scientific trial. Publish negative and inconclusive results
within the same agreed privacy boundary.

## Relation to the current public service

The public registry offers T01 Goal continuity, T02 causal state controls,
and T03 bounded preference learning and reuse. Its offline kit verifies signed
proofs; it does not run this succession protocol or execute D0/D1/D2 on the
evaluator's machine. See the
[availability matrix](scientist-verification-matrix.md),
[public downloads](./#downloads) and [new continuity challenge](../eval/#new-challenge).
