T07 · Researcher-authored challenge

Your experiment.
Your unseen tests.

Give DRENDA observations, not the answer. Freeze its learned program, then measure what survives new inputs, ablation and a fresh process.

01 · Observe02 · Freeze03 · Test04 · Remove & restore05 · Verify

What is actually tested

You choose the problem and its data. The frozen DRENDA typed-program synthesizer searches its existing bounded grammar using only your visible examples. Your category is metadata; your description is recorded, not interpreted as a natural-language instruction by this component.

The hidden answers stay with the service and separate Judge. Hidden inputs are released only after the learned artifact has been frozen. Baseline, learned, ablation, exact restore and restart all use the same held-out rows. No shell, user code or private Resident is exposed.

Input conventions, limits and interpretation
  • Use JSON numbers, strings, booleans, null or nested arrays. Objects inside input/output are not supported.
  • An input array is a list of arguments. To supply one sequence argument, wrap it in another array: [[…]].
  • Provide 2–64 visible examples and 2–64 hidden tests. No repeated inputs and no train/hidden overlap.
  • Up to 32 items per array, 512 characters per string, four nesting levels and absolute numeric values ≤ 1,000,000.
  • Available grammar includes arithmetic, comparisons, Boolean and bounded sequence/text operations. It is not an arbitrary software or unrestricted prose solver.
  • PASS requires every hidden row correct, a gain over the initial identity primitive, loss under artifact removal, exact restoration and fresh-process retention. PARTIAL, FAIL and infrastructure errors remain visible.
  • You supply the reference answers. Verification proves performance against them, not their mathematical truth. The server operator is trusted to maintain the seal; this is not secrecy against its administrator.

Design the challenge

Checking availability…

Version 1.1 prunes type-incompatible combinations before evaluation. Receipts separately count raw combinations, pruned combinations and interpreter evaluations (including failed or duplicate evaluations). Each challenge starts fresh.

DRENDA may use these inputs and answers to construct its program.

DRENDA never receives these answers. Keep inputs distinct from the visible observations.

Submit synthetic or non-sensitive research data only. The full proof contains your observations and reserved answers. It is accessible using your private receipt, not publicly by challenge ID.

Resume an experiment

Verify, do not just trust the badge

Download the proof and offline verifier. The verifier checks Ed25519, file hashes, exact artifact lineage, process receipts, freeze order and all five scores against your reserved answers.

python -m t07_modules.t07_verifier proof.zip \
  --trusted-key-sha256 f27be1d8450c1f3ce920a614ebfb306925a388333f243aabe5d9ccd184aaffe3

Requires Python 3.10+ and cryptography. Verify the key through a trusted channel; a proof carrying its own replacement key is not sufficient. A valid signature can also authenticate a FAIL — integrity and performance are separate.