T07 · Bring your own experiment
Choose the problem. Reserve the answers.
Create your own input/output challenge. DRENDA learns from visible observations; a separate Judge checks your sealed tests, ablation, exact restoration and fresh-process retention. Download a private, independently verifiable proof.
Design and run your own challenge →
Frozen bounded program grammar; structured data only. The description is recorded, not interpreted as an unrestricted natural-language request. T01–T06 remain available below.
DRENDA Public Evaluation Lab
Test the creation of an executable tool from experience, learning and reuse on new tasks, or Goal continuity. Remove an artifact, restore it and check retention in fresh processes. Inspect the separate Judge's decision and verify the signed proof yourself.
LIVE CHALLENGE ONLINEOFFLINE PROOF VERIFIERRESEARCH CASES
The live experiment
T06 · Does learning in text improve acquisition in numeric tasks?
DRENDA's bounded synthesizer builds M1, then reuses it in text capability M2. Validated reuse updates its search policy. On a separate numeric branch, it acquires N1 and attempts N2 under the same data and a 500-candidate search budget.
Compare with text experience → without it → remove only the learned policy → restore exactly → restart. The controls preserve the numeric parent and all observations. A larger 50,000-candidate diagnostic budget measures the additional search cost; it is explicitly not the matched-budget comparison.
Four composition forms were known during development. Each public trial creates new random study observations; hidden outputs are drawn after candidate freeze. The candidate sees no teacher formula or holdout labels. This measures transfer of a reuse policy between text and numeric search, not arbitrary cross-domain abstraction.
The search-order repair and protocol were externally engineered; the bounded DRENDA mechanism synthesizes the particular operators and updates its policy from validated experience. The earlier 16-example protocol failed and is preserved; the frozen public amendment uses 64 generic observations. No draw is retried to obtain PASS.
The first public R15 deployment trial was INVALID because the DAC-restricted supervisor could not read a disposable checkpoint under the service's 0077 file-creation mask. Its original signed INVALID is preserved. The transport repair changes permissions only inside a root-private disposable lease; the candidate, data, Judge and acceptance rules are unchanged.
A separate Judge replays artifacts without importing the candidate. Download raw synthetic observations, checkpoints, receipts and signed verdict; verify all controls offline. The private Resident is never accessed.
T05 · Compose a learned capability into a structurally different task
Family A maps a sequence to a Boolean. Family B maps a sequence and an independent context input to a Boolean, using a learned A artifact as an executable parent. This changes the input structure and dependency graph, not just names. The preserved DRENDA mechanisms receive observations, not a solution program or hidden answers.
Compare B without A → B with learned A → remove only A → restore the same hash → restart. The baseline runs two existing acquisition routes on the same 62 B observations and frozen limits. All B conditions use the same 256 held-out inputs. The page separates wrong answers, abstentions, and execution/dependency errors. Removing A makes its executable dependency absent; this is not 256 wrong answers.
This is a bounded composition experiment in existing representation languages, not arbitrary or distant cross-family transfer. New server-side draws rename symbols/fields and reorder a fixed exhaustive support. The post-freeze draw changes holdout order, not its support. Seven fresh candidate processes have a maximum of 90 seconds each; one attempt is retained, including FAIL or INVALID.
T05 accepts only a private challenge label, never code, seeds or a problem description. It is available only when confirmed by the registry. Public verification replays generated data artifacts independently; the learner runs only on the private server. No private mechanism source is downloaded.
Select T05 only when registered · Read the frozen T05 protocol and limits
T04 · Create an executable tool from observations
The preserved DRENDA synthesis mechanism receives 31 observed input/output transitions, with opaque names and symbols. It constructs a data-only finite-state tool, freezes it, then a new process executes it on 64 longer sequences generated after the freeze. No target machine, hidden answers or program is supplied to the learner.
Seven fresh processes cover no-tool baseline → construction → restart/reuse → tool ablation → exact restore → constant-output replacement → restore again. The separate Judge reconstructs the synthetic experiment and interprets the generated artifact independently. All draws and outcomes are preserved, including failures.
The existing search algorithm and representation language are engineering; the particular tool is synthesized by the unchanged DRENDA mechanism. This tests bounded tool construction, not arbitrary software creation. The no-tool baseline and ablation have no executable predictor; the constant-output replacement provides a behavioral control. This holdout stays within the same sequence family.
Select T04 when confirmed by the live registry · Read the tool-creation protocol
T03 · Learn, reuse, ablate, restore
The existing DRENDA experiment-selection learner receives eight generated demonstrations and synthesizes a ranking program. Only after that artifact is frozen are 48 new tasks generated, with new names, values, order and numeric outcomes instead of categorical labels. Every task is retained, including errors.
Five fresh processes measure: before study → learning → restart/reuse → learned-artifact ablation → exact restoration. The separate Judge recomputes the training and holdout scores. The result page shows the measurements; the proof reveals the synthetic dataset seeds, learned program and all five receipts for offline audit.
This is supervised learning of a preference over already-described experiments, using four existing structural features. The demonstrations are supplied by the evaluator protocol. The DRENDA learner, unchanged from its installed version, synthesizes the program; it is not supplied the teacher's coefficients or the holdout answers. Your text labels the run; it is not the problem being solved.
Run T03 below · Inspect the frozen learning protocol
Published deployment check, 25 September 2026: 40/48 new-task decisions after learning and restart; 5/48 with the learned artifact removed; 40/48 after exact restoration, on the same 48 tasks. The separate source baseline was 2/8, not the same dataset. This is an operator-assisted public run, not an independent reviewer result. Inspect this recorded result · Download its proof.
Transfer here means new tasks within the same four-feature grammar, not transfer to a different problem family. The proof includes all 48 outcomes, including the eight errors.
Choose a restart trial or the new causal state-control battery. Both execute the same frozen DRENDA state component with your own Goal. The second trial runs eight fresh processes: baseline creation/restart, then three perturbation/restore pairs. A separate Judge signs the result; the proof can be checked offline.
Remove the Goal
Remove only the active Goal and recompute the state commitment. Does resumption fail even with a consistent hash? Restore the original bytes: does the same Goal return?
Try a foreign Goal
Keep the checkpoint but change the challenge identity. Does the probe reject this mix-up without changing state? Does the original input still work?
Corrupt the state
Change the persisted revision without changing its commitment. Does the probe detect it? Does exact restoration recover the original state?
T02_CAUSAL_STATE_CONTROLS requires all three rejections and all three recoveries. A successful baseline alone cannot pass this contract. These are controlled state experiments, not learned-capability ablation or repair authored by DRENDA.
T01/T02 measure exact Goal-state continuity. The evaluator writes the Goal; the fixed probe records it; the Judge checks identity, content and state commitments after restart. T03 measures the learning protocol above. Inspect each contract and its authorship in the protocol.
Run and observe
Submit an evaluator-created Goal to the isolated public worker and follow the queue, process restart and separate Judge receipts.
Check independently
Download the signed proof and follow the offline verification guide. Inspect bundle hashes, event chain and the exact acceptance predicate.
Submit no passwords, personal data, secrets, or confidential research. The plain-text Goal is temporarily processed by the service to run this fixed test; its text is not included in the downloadable proof.
The wider DRENDA research program
Scientists can examine the live public challenge, inspect documented internal research cases and design supervised post-freeze trials. Each capability is evaluated separately under a frozen budget, hidden cases, causal controls and a Judge distinct from the candidate.
| Research question | Decisive test | Evidence route |
|---|---|---|
| Does a Goal survive process restart and causal controls? | Fresh processes, Goal ablation, foreign input, corruption rejection and exact-byte restore. | Run either public contract |
| Can one acquired capability enable the next? | Equal-budget D0/D1 comparison, post-freeze holdout, D1/D2 ablation and restore, restart and transfer. | Inspect the recorded D0 → D1 → D2 case |
| Can DRENDA sustain grounded dialogue? | Evaluator-written multi-turn questions, distractor Goals, restart, ambiguity controls and scored answers. | Independent protocol |
| Does learned text reuse reduce numeric acquisition cost? | Matched-budget policy/control, policy-only ablation, exact restore, post-freeze holdout and fresh-process retention. | Inspect and run T06 |
| Can it learn and transfer from experience? | Eight demonstrations, 48 new tasks in the same structural feature space, scratch baseline, learned-artifact ablation, restore and fresh-process reuse. | Run public T03 |
| Can it construct an executable tool from observations? | New synthetic observations, a synthesized finite-state tool, post-freeze hidden sequences, independent interpretation, ablation, sham, restore and restart. | Inspect and run T04 |
| Can it author and validate an executable repair? | Frozen fault, concealed fix, DRENDA-authored candidate, hidden tests, rollback, restart and same-Goal continuation. | Independent protocol |
| Can it investigate new science or mathematics? | Post-freeze question, explicit predictions or proof obligations, falsification and independent review. | Independent protocol |
| Can it operate over long periods without intervention? | Hours-to-days run, intervention ledger, resource ceiling, incidents and independently observed checkpoints. | Independent protocol |
The links distinguish the public live challenge, an internal research record and protocols for further independent trials. See independent evaluation requirements to arrange a scoped scientific test.
Run the registered challenge
Choose the contract below. The registry must confirm that it is available. Input is plain text, limited to 500 characters / 2,000 UTF-8 bytes.
Create a NEW trial, then verify its own proof
- Wait for the live registry confirmation below. Write your own non-sensitive Goal and a fresh arbitrary label, such as
trial-your-chosen-label. Do not reuse the sample as evidence of a new run. - Keep your exact text and trial label locally. Confirm the scope checkbox, then press Submit challenge once. A new receipt gives a new
ch_…challenge ID; save it immediately. - Watch the worker/Judge status. If automatic polling stops, use Refresh status; do not submit a duplicate. After reload, use Recover an existing result below with your saved ID. This makes read-only requests, not another submission.
- When terminal, download that challenge's ZIP, not a historical sample. Verify it offline and match the returned ID to your receipt. Keep the ZIP hash, exact verifier output, environment and any FAIL/INVALID result.
T01/T02 preserve your statement as Goal data. T03/T04 use it only as a private run label and generate new synthetic tasks. T05 also accepts only a label and randomizes names/order over fixed exhaustive supports. T06 uses a label and fresh random observations in four fixed composition forms. None interprets the statement as an open-ended task. QUEUED, a timeout or an unavailable proof is not PASS.
Level 3: verify on a second machine
Browse all public proof and evidence downloads. Existing signed proofs are historical examples; use the new challenge ID above for your own trial.
The kit contains the offline verifier, the public Judge key, a previously generated sample proof, and the fixed acceptance rules. Its current status is prepared, not independently run on a second machine. Download it, verify the pinned key fingerprint through a separate trusted channel, and run the proof verifier on a separately administered computer.
The evidence-pack command python VERIFY_EVIDENCE.py does not verify a signed challenge ZIP. Use python -m level2.verify_proof challenge-proof.zip judge-public.pem from the extracted kit. Read verified and verdict separately: an authentic failed result is not a passed challenge.
Evaluator-created challenge ch_0a78fee6018ddd7933efe90763a17ce16464 returned PASS / GOAL_CONTINUITY_VERIFIED on 2026-09-23. Its public proof passed the offline verifier against the published Judge key: signature, hashes, event chain and acceptance predicate. Proof SHA-256: 29444f6a0b084daf8bd4206cb70dd7cbfb43b18bbc151937fd693949500d8e6e.
Evaluation boundary
| Candidate | T01/T02: frozen r26 cognitive-state component. T03: existing criterion-3 learner. T04: preserved cognitive-primitive and Omega synthesis mechanisms. T05: acquired-capability composition. T06: bounded text-to-numeric reuse-policy transfer, with an explicit externally engineered search-order repair. Wrappers transport typed data; distinct frozen bundle hashes are published in the registry. |
|---|---|
| Worker | Runs a constrained test in a disposable sandbox; no candidate network, production Resident, or user-supplied executable. |
| Judge | A distinct attesting process with a separate Ed25519 key. |
| Authority | Fixed gateway policy; a challenge result does not promote code or grant system access. |
| Independent check | Offline verifier validates bundle membership, hashes, event chain, acceptance predicate, and Judge signature. |
Novos testes abertos
T08 · Criação de ferramentas causais · T09 · Aprendizagem pareada limitada
Envie um novo desafio pela API isolada. Verdicts e falhas são preservados; a disponibilidade do experimento não é um PASS geral de capacidade.