Research access · reproducible evaluation · independent testing

Test the system. Trace the result.

Start with an executable public challenge. For a deeper evaluation, agree on a frozen build, new tasks and measurable acceptance criteria. An initial review does not require the private source repository or the creator's memory and data.

New challenge: step-by-step · All public proof and evidence downloads · Available tests versus proposed capabilities

Choose the evidence you want to examine

Available now · public

Bring a fresh Goal

The registered continuity challenge runs in an isolated worker, restarts a frozen DRENDA state component and returns a separately signed result. Read the contract before submitting your non-sensitive Goal.

Open the live test
Available now · inspect

Verify and challenge the record

Inspect the technical report and causal succession case. Download proof-verification tools. Checking an existing proof is distinct from independently rerunning its underlying experiment.

D0 → D1 → D2 case study
Offline proof-verification kit
By arrangement · scoped

Design the next evaluation

Propose post-freeze tasks and an independently administered environment. Access, build scope, resources and confidentiality must be agreed before a broader run. This page does not provision access to the private Resident.

Trial requirements
A proposed protocol, not a completed result

Make the next claim earn its evidence

Select one capability and freeze its measurement before execution. The public continuity contract remains the only registered live test; the controls below describe the requirements for a separately arranged, broader trial.

Declare the evaluator's role. Owner-authorized technical review may be assisted by an AI tool. Label that work as AI-assisted technical evaluation, record the tool/version and human interventions, and preserve its artifacts. It is not a human-expert endorsement or independent external validation. Access to hidden answers must be disclosed; a shared assistant context does not establish blind judging.

  1. Freeze: record the exact build and state hashes, tool permissions, machine, task family, time and memory limits, and acceptance criteria.
  2. Keep evaluation independent: the evaluator prepares fresh tasks after the freeze. Hidden answers, expected repairs and private Judge inputs stay out of the candidate's environment.
  3. Measure behavior: preserve inputs, actual outputs, failures, resource use and every human intervention. Identify whether an artifact came from DRENDA, engineering code or an external source.
  4. Test a claimed improvement: compare baseline and candidate under the same budget. Remove only the claimed learned artifact, repeat the task, then restore the exact artifact and repeat again.
  5. Test retention and reuse: restart in a fresh process, resume the same Goal, and try an unseen related task. Record a failed or inconclusive transfer without changing the criterion.
  6. Deliver an inspectable result: publish an agreed, sanitized proof bundle and report the independent evaluator's conclusion. A signature establishes provenance; it does not replace a sound scientific test.

For cumulative succession, use the D0 → D1 → D2 independent trial protocol. The current public case is a summary, not a downloadable executable reproduction. A new sanitized frozen package and evaluator-controlled run require separate review and agreement; this page does not grant access to private materials.

Propose a concrete evaluation

Include your institution or team, the capability you want to test, your success and failure criteria, available hardware, and whether you want to verify an existing proof or execute a new trial. Do not send confidential datasets or credentials in an initial message.