Skip to main content

Form an evaluation

Merion treats evaluation engineering as a governed compilation process—not as automatic trace conversion.

1. Register evidence governance

Every episode must reference an active data-use policy that permits evaluation. Register the policy before analysis:

2. Analyze a cohort

The cohort records the sampling frame, inclusion and exclusion rules, schema inventory, coverage gaps, and proposed clusters.

3. Qualify a task candidate

Qualification explains economic importance, recurrence, capability uncertainty, representativeness, reproducibility, verifiability, non-triviality, rights/privacy, and outcome quality.

4. Draft four independent specifications

An evaluation proposal contains separately digest-pinned specifications:
  • TaskSpec: realistic instruction, initial state, allowed actions, and expected end state.
  • HarnessSpec: prompt, tools, permissions, budgets, and stop rules.
  • EvalEnvironmentSpec: reset behavior, fidelity evidence, omissions, dependencies, and secrets boundary.
  • VerifierSpec: capability axes, checks, calibration evidence, exploit hypotheses, and disagreement policy.

5. Approve and audit

Each current specification digest requires an independent human approval. Changing a specification invalidates its approval. After approval:
The audit checks resetability, data rights, exposure compatibility, solution leakage, environment fidelity, verifier independence, reproducibility, and human solvability.

6. Pin the experiment

This validates the exact task, environment, harness, and verifier boundary. It does not execute trials or train a model yet.
Last modified on August 21, 2026