Both configurations run the 19-case Hyperion development pack. Both are single-arm and descriptive: they estimate one generation system's exercise success rate with a case-clustered bootstrap. Neither declares a contrast, because there is only one arm to compare.
benchmark.smoke.yaml is executable now. It runs the deterministic fixture generator over all 19
cases and exercises dataset loading, planning, generation, evidence capture, and release-ready
attempt records:
bun run cli validate studies/hyperion-development/benchmark.smoke.yaml
bun run cli run studies/hyperion-development/benchmark.smoke.yaml --id smoke-1Generation succeeds 19/19. The configuration declares no evaluator, so it does not exercise
evaluation — that is a separate command with its own configuration
(process evaluators). A clean run never enters a resumable
state either, so it does not exercise resume; re-running with the same ID fails with EEXIST.
benchmark.full-stack.template.yaml defines the live generation system through the external
Artemis production-API adapter. It requires a normally configured full Artemis deployment with
core, localci, localvc, Hyperion exercise generation, a model endpoint, build agents, a
dedicated benchmark course, and environment-provided editor credentials. Before a live run:
export EXGEN_ARTEMIS_ADAPTER_ID=artemis # must exactly match this system's systems[].id
export ARTEMIS_BENCHMARK_USERNAME=...
export ARTEMIS_BENCHMARK_PASSWORD=...- replace every placeholder revision, URL, course ID, and fixture default; use a separate system entry (and matching adapter ID) for every arm;
- pass credentials through the declared environment-variable mapping, never the YAML file;
- record the Artemis, database migration, LocalVC, LocalCI, build-agent, sandbox, model-serving, prompt, and effective configuration digests;
- restore a verified logical baseline and give every attempt fresh exercise/repository state;
- freeze routing, toolchain, budget, evaluator suite, retry policy, concurrency, and seeds;
- run
exgen validateand archive the resolvedexgen plan --json; - evaluate every frozen candidate with an independent pinned verifier; and
- label results descriptive, because this public corpus is a development set.
The production API flow, evidence layers, and formal acceptance gates are specified in
docs/ARTEMIS-INTEGRATION.md. The failure, retry, concurrency,
and registration rules that govern the run are in
docs/METHODOLOGY.md.
Steps 3 to 5 remain manual. The configuration records declared deployment deviations, but no code
resolves the Artemis build and profile contents, restores a verified baseline, or compares a
generator-visible fixture across attempts. This template stays single-arm because the limitations
in the integration design still prevent a
defensible comparison. A second arm can select another effort_profile on the same deployment, but
configuration alone does not make the contrast valid.
19 cases at one replicate estimate a rate to roughly ±20 percentage points, and cannot support a
comparison of anything smaller than about a third of the outcome scale. Declare the smallest
meaningful effect before running, and check it against
docs/METHODOLOGY.md.
19 cases × 3 replicates = 57 attempts. At the enforced execution.concurrency: 1 and declared
wall_time_ms: 3600000, the worst case is
57 hours of continuous serial execution. The expected
wall clock depends entirely on the deployment and has not been measured here, so the first task
before scheduling a live run is a pilot of a few attempts to obtain a median duration; a run that is
scheduled on the worst case alone will be abandoned halfway.
The template declares no stopping rule, and one must be added to the registration before use.
docs/ARTEMIS-INTEGRATION.md requires a stopping rule in the frozen registration, and a run
abandoned without one produces unstarted attempts that stay in the denominator and block a
confirmatory release.
The template declares only the wall-time budget. Add call, token, or cost budgets only when every
terminal job reports accountingState: COMPLETE. The adapter waits while the state is PENDING and
treats INCOMPLETE as a permanent accounting gap. Artemis's effective internal token and quota
policy remains part of the system manifest.
Planner and executor model assignments are explicit factors of the generation system. A future strong-planner / weak-executor comparison should be a predeclared 2×2 factorial with fresh paired attempts, a fixed orchestration and evaluator, role-level call/cost evidence, and interaction analysis. Do not silently route, retry, or fall back between models.
Each configuration uses two distinct seeds so the schedule and the resampling draw are independent,
and so the smoke and live studies do not share a draw. Each is the first eight hex digits of
sha256("exgen-bench/<config id>/<purpose>") read as an unsigned 32-bit integer:
| Configuration | trials.base_seed |
analysis.bootstrap_seed |
|---|---|---|
hyperion-development-smoke |
497506175 | 2946802029 |
hyperion-development-live |
3338702608 | 150218520 |