# DRENDA public evaluation protocol

## Available public tests

The public route `/drenda/eval/` is a static interface to a separate bounded
challenge service on the VPS. The baseline contract is:
`T01_T14_PERSISTENT_GOAL_CONTINUITY` (v2.0.0). It checks whether one
evaluator-supplied Goal keeps its identity and content commitment when the
frozen r26 DRENDA cognitive-state component is rehydrated in a fresh process.
The public test registry is the authority for its exact contract and current
bundle hash.

## T06 cross-domain search policy

`T06_CROSS_DOMAIN_SEARCH_POLICY` contract v1.0.0 is additive. The enclosing
queue/proof envelope stays v2.0.0; neither version is inferred from a PASS.
Only a non-sensitive `goal_text` run label is accepted at `/v2/challenges`.
No caller-selected seed, teacher, code, path or tool is accepted.

The frozen candidate source contains an externally engineered search-order repair.
The particular operators and reuse-policy update are produced by the bounded
DRENDA mechanism. The protocol supplies study observations, not a target AST.
This is not a spontaneous live Resident repair or general new representation.

M1 learns text concatenation; M2 composes a transformed token with an acquired
text parent. Independently scored M2 reuse changes `prefer_accumulated_reuse`.
N1 acquires numeric absolute difference in BOTH branches. The policy branch
retains text experience; scratch lacks it. N2 uses 64 fresh ordinary random
numeric observations from one of four development-known forms. The candidate
is not shown the form, seed, hidden labels or expected expression. Candidate
artifacts are frozen before each 64-case holdout draw. All generated data,
failures and error rows are retained. There is one scientific draw, no retries.

Three N2 branches use identical observations and a 500-candidate budget:
policy, scratch and policy-only ablation. Scratch/ablated diagnostic branches
then use 50,000 candidates solely to measure acquisition cost, not as equal
budget competitors. Exact policy restoration must recover the same candidate
and cost. A new process must retain its 64/64 numeric predictions without
training or state changes. No sham result is fabricated or inferred.

PASS requires all M1/M2/N1/N2 holdouts 64/64, policy acquisition within 500,
no executable scratch/ablation candidate within 500, correct diagnostic
controls at strictly greater cost, exact restoration, and fresh-process
retention. Isolation, identity, raw hashes, causal state checks, or freeze
ordering failures are INVALID; a valid run lacking gain is FAIL. Fixed rules
are in the public verifier; candidate code is not included.

The earlier 16-observation development protocol FAILED: a spurious control
fit the study but scored 54/64 hidden cases. It is not reclassified. The public
64-observation amendment was frozen separately and evaluated on 20 random
draws across the known four forms, with every draw retained. This is bounded
repeatability, not a universal generalization claim.

21 fresh networkless candidate processes, independent phase scorers and a
separate Ed25519 Judge retain raw synthetic requests, outputs, program data,
checkpoint snapshots, process identities and metrics. Resource ceiling:
600 seconds candidate wall budget, 1 GiB, one queue worker, 720-second lease.
The verifier recomputes observations from disclosed seeds, artifact semantics
and causal controls; it does not run the private engine. Host receipts are
server attestations, not independent hardware attestation. Independent
execution on a separately administered host remains a different experiment.

The label and lease secret are omitted; only commitments appear. No Resident,
private memory, source, credentials, shell or promotion endpoint is public.
The protocol's old FAILs remain documented even when a new draw passes.

## T05 bounded cross-family composition

`T05_CROSS_FAMILY_COMPOSITION_TRANSFER` v2.0.0 is an additive, private-server
execution path. It is selectable only when its frozen private bundle is configured
and the live registry confirms availability. Preparation of this page/kit does
not assert deployment or a verified Linux run.

The V2 protocol commitment is
`cd9cae245619dfdaf9f61ebcf1bc40fde284374fd27cad55847e98664b66e7fd`.
Family A is sequence → Boolean, represented by an acquired finite-state artifact.
Family B is sequence + independent context → Boolean, represented by a conditional
two-feature rule referring to A by its exact hash. This is a structural input and
dependency change, but still composition within existing representations. It is
not arbitrary, distant or general transfer, and does not measure live Resident/N15.

Frozen supports and counts: A learns from all 63 binary sequences of lengths 0–5
and is held out on 64 sequences of length 6. B learns from all 62 sequence/context
pairs of lengths 0–4 and is held out on 256 pairs of length 7. Server-side random
draws change opaque names, field names and order only. The draw after artifact
freeze permutes the same exhaustive holdout; it does not create unseen support.
No draw is selected, retried or retuned to obtain PASS.

Seven fresh candidate processes execute A learning, B learning without A,
B learning with A, A holdout, baseline B holdout, causal controls, and restart.
Both B acquisition conditions run the existing open-primitive route and existing
Omega raw-feature route with the same observations and configured limits. The
parent-requiring route alone is not the baseline. No candidate from both routes
is recorded as abstention/unavailable, not as proven inability of every DRENDA
mechanism. Budget: 90 seconds/process (630 seconds maximum candidate wall time),
128 observation ceiling, three DFA states, feature arity two, at most one parent,
two outer routes/acquisition, and zero holdout training updates. Equal configured
budgets do not assert equal actual CPU work. One worker and one Judge attempt;
the queue lease is 720 seconds. Operational failure may be INVALID with no proof.

The candidate receives only training observations or held-out inputs and the
allowed generated parent artifacts. It sees no teacher formula, holdout labels,
seeds, submitted label, Judge key or supervisor evidence mount. Public input is
exactly `{"challenge_label":"a non-sensitive trial label"}` at
`POST /drenda/eval/api/v2/cross-family/challenges`; no user code or seed is accepted.
Label limits are 500 characters / 2,000 UTF-8 bytes; control characters are denied.

The controls remove only A, restore its exact bytes/hash and restart in a new
process on the same B inputs. Artifact ablation creates a missing executable
parent dependency, not incorrect Boolean answers. Scores explicitly distinguish
correct, wrong, abstained and execution-error rows. A PASS requires 64/64 A and
256/256 B correct, at least 0.25 accuracy gain over baseline, at least 0.25 drop
on ablation, and exact restored/restarted outputs and artifact identities.
Those thresholds and the learner are unchanged from the preregistered V2.

Candidate subprocesses have private tmpfs work areas, no network and no access
to the private Resident. Artifacts cross stdout; only the supervisor writes its
host-side persistence. The strict runner has no local fallback and requires
verified Linux service restrictions. Windows engineering receipts remain INVALID;
their original CRLF bytes/hashes are preserved. Local unit fixtures simulating
Linux flags are not Linux isolation evidence.

After execution, a separate Judge independently interprets the generated data
artifacts and reconstructs synthetic observations/holdout from disclosed seeds.
The public kit includes this data-artifact interpreter and fixed rules, not the
private learner, Candidate wrapper or worker. Offline proof replay is not an
independent rerun of the engine. The existing sample/key are retained unchanged:
the historical sample proves T01 only. Preserve every FAIL, INVALID and unavailable
receipt; neither successful signature verification nor mechanical dependency
alone establishes cognitive transfer. A positive run is bounded evidence of
composition using an acquired capability, not a broad transfer claim.

T05 limits are scoped: label request 4,096 bytes; authenticated private worker
result 2 MiB; proof ZIP/member 2 MiB and total uncompressed 4 MiB; JSON depth 32,
100,000 nodes, finite numbers only, duplicate keys denied. Legacy test limits
remain unchanged. Public outputs exclude source code and submitted labels.

## T04 executable tool creation

`T04_BOUNDED_TOOL_CREATION` uses unchanged DRENDA cognitive-primitive and
Omega synthesis mechanisms in a private frozen candidate bundle. It accepts
only fixed, bounded data operations, never shell commands or submitted code.
The public text is a run label, not an open-ended task.

The evaluator uniformly samples a reachable two-state binary machine with
opposite Boolean state outputs, supplies all 31 sequences of lengths 0–4,
and conceals the machine, seeds and hidden outputs from the candidate. Names
and symbols are opaque. DRENDA synthesizes a finite-state data artifact using
its existing search language; no new interpreter or operator is claimed.

Only after that artifact is frozen, a separate seed generates 64 distinct
sequences of lengths 6–24. No outcome-based filtering or retry selects a
successful machine. Seven fresh, networkless processes measure the no-tool
baseline, construction, restart/reuse, artifact ablation, exact restoration,
a correctly rehashed constant-false replacement, and exact restoration again.
Only the tool checkpoint crosses candidate processes.

The baseline and artifact removal have no executable predictor. They show
availability and dependency, not a competing cognitive strategy. The sham
replacement tests behavioral content. The frozen predicate requires 64/64
correct hidden outputs, at least eight more correct outputs than the sham,
and identical restored/restarted execution. A heavily imbalanced draw can
fail the gain threshold even with correct execution; that FAIL is retained.

The separate Judge reconstructs the complete synthetic dataset and interprets
the artifact independently, without importing the DRENDA learner. The signed
proof discloses the synthetic seeds, generated artifact and process receipts
after execution, but not private source, Resident memory or the submitted text.
The offline verifier recomputes the predicate and verifies the signature.
Local Windows engineering tests are explicitly not a verified Linux sandbox
and cannot receive a public PASS under the strict isolation requirement.

The trial shows bounded tool synthesis and use on longer sequences within the
same family. It does not establish transfer between different problem families.
Available contracts and exact frozen bundle hashes are authoritative in the
live registry; an unavailable T04 option must remain disabled.

## T03 learning and reuse

`T03_LEARNED_EXPERIMENT_PREFERENCE` runs the existing DRENDA criterion-3
transfer engine, byte-identical to the installed learner. A new fixed wrapper
only transports typed demonstrations and tasks and persists the learner's
output. No private Resident, memories or source code are downloadable.

The experiment is supervised program synthesis in an existing four-feature
space: partition balance, expected posterior size, distinct outcomes and cost.
The evaluator privately samples a preference objective, supplies eight example
choices, then generates 48 new tasks after the learned artifact is frozen.
Names, values and probe order change; source outcomes are categorical and
transfer outcomes are numeric. No task is selected based on candidate success.
Costs are sampled without replacement to avoid indistinguishable source options.
All errors are retained. The teacher's coefficients, seeds and held-out labels
are not in candidate inputs or mounted files. They become reconstructible from
the proof only after execution. This is not raw semantic interpretation or
autonomous curriculum creation; it tests the existing learner on new experience.

Five fresh, networkless processes:

1. **Before study:** scratch policy chooses among the eight source experiments
   without seeing the demonstrated decisions. At least one error must occur.
2. **Learn:** DRENDA synthesizes a program from those eight demonstrations.
   The program and its training commitment are frozen; the process exits.
3. **Restart/reuse:** a fresh process reads that artifact and selects among
   the 48 new tasks. No training runs in this process.
4. **Ablation:** only the learned artifact is removed. With the exact same
   tasks, wrapper and learner, the scratch policy is measured again.
5. **Restore:** the exact artifact bytes are restored and a fifth process must
   recover exactly the learned decisions.

The separate Judge independently implements the preference score. It checks
all input/output hashes, source fit, post-freeze ordering, program/checkpoint
identity and restore equality. Frozen acceptance: at least 24/48 correct held-out
choices and at least eight more than scratch, with all causal controls passing.
Budget: at most 45 seconds/process, 225 seconds of candidate wall time; one job
at a time. Same tasks and one decision per task in both transfer conditions.
The public status exposes the measured counts, not only a PASS badge.

The proof contains the newly learned bounded program, synthetic dataset seeds
and five process receipts. It contains neither the submitted text nor private
memory/code. Its verifier reconstructs the complete synthetic training/holdout
and recomputes the scores without executing the learner. A signature attests
the server's recorded run; independent rerunning of the candidate on a separately
administered host is a distinct evaluation. A single PASS is evidence for this
run, not a guarantee of success on every fresh draw. Preserve FAIL and INVALID.

The mechanism, wrappers, data generator and Judge are engineering; the program
parameters are synthesized by the unchanged DRENDA learner from the examples.
T03 is a fresh execution, not a replay of the August internal result.

## Causal state controls (T02)

`T02_CAUSAL_STATE_CONTROLS` v2.0.0 uses the **same unchanged frozen bundle**.
It runs eight separate candidate processes: create, baseline restart, then
three negative-control/restoration pairs:

1. GOAL_ABLATION: remove the active Goal and recompute the state commitment.
   Resumption must reject this otherwise internally hash-consistent state.
2. FOREIGN_GOAL_INPUT: leave the checkpoint untouched and submit a different
   challenge identity. It must be rejected without modifying the checkpoint.
3. STATE_TAMPER: change the saved revision without updating its commitment.
   The damaged state must be rejected.

After each control, the worker restores **exactly the original checkpoint
bytes**, not a freshly reconstructed Goal, and starts another process. Its
observation must equal the successful baseline restart. Every negative must
exit with the frozen probe's controlled rejection (not a sandbox crash).
The separate Judge and offline verifier require all three rejections,
unchanged state on rejection, all exact-byte restorations, output commitments
and all three recovery observations. The new test cannot pass on continuity
alone. Budget: 20 seconds per process, 160 seconds total, one queued job at
a time. The API registry determines which contracts are actually deployed.

The perturbations and Judge are externally engineered test infrastructure. The
candidate component is not retrained or modified. This measures storage and
validation of Goal state; it is **not** ablation of a learned capability,
semantic understanding or autonomous repair. Signed receipts
bind what this worker reported, not an independent attestation of the host.
The public proof contains hashes and observations, not the checkpoint or
the submitted text. Use the updated verifier kit for T02; old T01 proofs
remain valid with the updated verifier.

An input is plain text only, limited to 500 characters / 2,000 UTF-8 bytes.
Do not submit secrets, personal data, or confidential material. The text is
temporarily processed to run the challenge and is not included in the proof
bundle. The service rejects unregistered tests, arbitrary code, uploads, shell
commands, or caller-selected tools.

The worker executes a constrained test in a disposable sandbox. It cannot
reach the private Resident or production memory. A distinct Judge process
attests the outcome. A PASS applies only to this registered continuity test;
other capabilities require their own registered tests and evidence.

## Create a new challenge, not a sample replay

1. Open [the public form](./#new-challenge) and wait for its live registry
   confirmation. Save the test version and frozen bundle hash from the
   [registry](https://trinh3.com.br/drenda/eval/api/v2/tests).
2. Write your own non-sensitive statement and a fresh arbitrary trial label.
   Keep the exact submitted text locally. The browser trims leading/trailing
   whitespace. The probe preserves text as data; it does not answer a question
   or perform an activity described in that text.
3. Acknowledge the scope and submit once. A valid HTTP 202 acceptance receipt
   identifies a new `challenge_id`. Save that ID before leaving the page.
   Neither acceptance nor QUEUED means the test passed.
4. Follow status and receipts. Automatic polling is bounded; use Refresh
   status after it stops. A timeout is an observation failure, not a verdict.
   Do not repeatedly submit to obtain a desired outcome.
5. At terminal PASS, FAIL or INVALID, download the available proof. Preserve
   the status response and ZIP even when the outcome is negative. If no proof
   is available, report that limitation; do not replace it with the sample.
6. Verify the ZIP using the commands below. Match its challenge ID to the new
   receipt, its test/version and candidate commitment to the frozen registry,
   and its key fingerprint to a separately trusted record. Preserve exact
   output and exit status, proof hash, date, Python/dependency versions and
   any human intervention.

If the page was reloaded, replace `YOUR_SAVED_CHALLENGE_ID` in these read-only
addresses with the exact ID from your acceptance receipt:

```text
https://trinh3.com.br/drenda/eval/api/v2/challenges/YOUR_SAVED_CHALLENGE_ID
https://trinh3.com.br/drenda/eval/api/v2/challenges/YOUR_SAVED_CHALLENGE_ID/proof.zip
```

If a submission loses its network connection before returning an ID, its
acceptance is unknown. Do not assume it was rejected or automatically retry;
keep the local record and report the missing receipt. HTTP 422 means the
request was rejected; a rate/queue limit must be respected. No identity,
credential, upload or alternative endpoint is needed to bypass a limit.

A new statement in T01/T02 is a new execution of the same registered continuity test,
not an unseen semantic reasoning problem. It cannot support a D0/D1/D2,
language, learning or autonomous-repair conclusion. See the
[availability matrix](../evidence/scientist-verification-matrix.md) and the
[separate succession protocol](../evidence/independent-succession-protocol.md).

## Level 3 independent verification

The downloadable second-machine kit contains an offline proof verifier, the
public Judge key, acceptance rules, and a sample signed proof. Its current
state is prepared; an independent second-machine run is pending. Before
relying on the Judge signature, verify the key fingerprint through a separate
trusted channel. Preserve both successful and failed verification results.

Use the [offline verification guide](../verify/#signed-proof) for downloads,
key-pinning checks and exact commands. After extracting the kit and installing
its pinned dependency, verify the included historical sample:

```text
python -m level2.verify_proof sample-proof.zip judge-public.pem
```

For a fresh challenge, download its proof, retain its challenge ID and save the
ZIP as `challenge-proof.zip` in the extracted kit directory:

```text
python -m level2.verify_proof challenge-proof.zip judge-public.pem
```

Match the returned `challenge_id` to the new trial. `verified: true` and
`verdict: PASS` have different meanings: the former accepts the proof's
signature/integrity and frozen predicate; the latter is the bounded test
outcome. An authentic FAIL/INVALID is not a successful challenge. The separate
`VERIFY_EVIDENCE.py` checks only the public evidence documents, not a signed
challenge ZIP. Offline proof verification also does not execute the candidate
on the evaluator's machine. T03 evaluates bounded learning through the separately
specified protocol above; T01/T02 do not evaluate learning.

## Public API

- `GET /drenda/eval/api/health`
- `GET /drenda/eval/api/v2/tests`
- `GET /drenda/eval/api/v2/verifier-key`
- `POST /drenda/eval/api/v2/challenges`
- `GET /drenda/eval/api/v2/challenges/{challenge_id}`
- `GET /drenda/eval/api/v2/challenges/{challenge_id}/proof.zip`

Example request body:

```json
{
  "test_id": "T01_T14_PERSISTENT_GOAL_CONTINUITY",
  "goal_text": "After restarting, continue this same evaluator Goal."
}
```

The interface and API are intentionally not a general-purpose DRENDA chat.
Candidate, Judge, and Authority remain separate. The private Resident,
Telegram, TRI, local memory, shell, filesystem, and promotion controls are not
exposed by this public route.
