blockedAPG3 stopped at its mandatory harness preflight. The current internal-agent harness could create fresh agents and return bounded results, but it did not expose a skill-invocation event or a supported per-agent skill profile. An overlapping delegation and assignment-composition skill collection was active, so baseline and candidate conditions could not represent the accepted target profile without global mutation, an adapter, or an invented substitute.
Accepted ADR 0002 requires APG3 to stop under those conditions. No evaluator was assigned, no evaluation material was created, no baseline was run, and no candidate was authored.
APG3 was authorized to evaluate whether ordinary Codex assignment composition
has a material deficit and, only if a frozen baseline-adequacy rule demonstrates
that deficit, author and evaluate at most one reversible
composing-bounded-worker-assignments candidate. The candidate could apply only
after delegation was already authorized and independently selected.
The authorization allowed APG3 to narrow or stop for evidence or environment limitations. It did not authorize a harness adapter, evaluation framework, plugin mutation, second skill, validator, runtime, dependency, taxonomy, publication decision, or licensing decision.
The bounded environment class was the current Codex desktop internal-agent harness with the repository instructions and installed skill-discovery surface available on 2026-07-18. The candidate leaf was absent. An overlapping external skill collection was enabled in the manager environment.
The preflight produced this capability disposition:
| Required property | Result |
|---|---|
| Fresh agent contexts | Available, but they inherited the same skill and plugin surface. |
| Bounded agent-harness returns | Available. |
| Observable skill-invocation event | Unavailable on the exposed harness event and tool surface. |
| Candidate available without forced invocation | No supported per-agent availability control was exposed. |
| Deliberate candidate loading | No supported internal-agent loading control was exposed. |
| Identical baseline and candidate environments | Not credibly maintainable without the missing controls. |
| Overlapping-skill isolation | No supported isolated internal-agent profile was available. |
| Sealed-input and downstream execution stages | Not reached after the mandatory stop. |
No evaluation environment was frozen because the hard preflight failed. An available command-line profile flag did not supply the missing internal-agent profile, invocation event, or treatment controls, and APG3 was not authorized to construct an adapter or weaker proxy.
The accepted contract required eligible families for small read-only work, isolated-write work, contract-sensitive work, disjoint selected worker scopes, and partial or needs-context outcomes. It also required non-trigger families for trivial direct work, unauthorized or unselected delegation, unresolved authority or design, coupled ownership, an already adequate assignment, and external prompts, scheduling, monitoring, managed reports, or manager-runtime requests.
These families remained authorization requirements only. No case text was created and no family was sampled because evaluation-role assignment occurs after the failed harness gate.
The accepted method would have separated unforced activation from deliberate composition, frozen distinct calibration and sealed confirmation evidence, bounded all samples and reserves, and required actual matched worker execution for downstream-compliance claims. The authoring gate would have required either a recurrent named binary failure or qualifying deficits across at least two anchored quality dimensions, or a qualifying manager-acceptability failure across a frozen share of eligible evidence. A lone outlier, minor correction, verbosity difference, or heading count could not qualify.
That method was not instantiated. No plan, rubric, baseline-adequacy rule,
reviewer contract, calibration bundle, sealed bundle, reserve bundle, sample
minimum, sample maximum, or execution subset was frozen. The executed sample
count and reserve use were both zero. The baseline rule was not applied, so this
outcome is not baseline-adequate-no-skill.
| Stage | Result |
|---|---|
| No-skill calibration | Not run. |
| Authoring decision | Not permitted because preflight did not establish a valid experiment. |
| Candidate calibration | Not applicable. |
| Activation | Not run; invocation was not observable. |
| Sealed confirmation | Not applicable. |
| Confirmation reviewers | Not assigned. |
| Reserve expansion | Not used. |
| Downstream execution | Not run; no downstream-compliance claim is made. |
| Manager acceptability | Not scored. |
| Binary evaluation safety | Not evaluated because no samples ran. |
The repository-level stop was safe: no candidate directory, empty placeholder, evaluation framework, fixture framework, validator, runtime, registry, scheduler, adapter, dependency, taxonomy, publication artifact, or licensing decision was added.
No candidate was created, so no candidate hash, correction, freeze, rejection
snapshot, adoption, or rollback mutation exists. The active skill tree remains
without a SKILL.md. The terminal skill disposition is not authored because
APG3 was blocked before evaluation.
Credible re-evaluation requires a supported internal-agent profile that removes the overlapping bootstrap and assignment-composition skills in both conditions, exposes an observable skill-invocation event, controls unforced availability and deliberate loading, and otherwise preserves model, harness, instructions, tools, and skill discovery. Providing or selecting such an environment is an external roadmap decision; APG3 did not create it.
External review is requested to accept the blocked disposition or authorize a separate roadmap decision about a supported evaluation environment. No later implementation phase is authorized by this result.
No claim is made about ordinary assignment adequacy, candidate quality, activation safety, downstream compliance, or whether a future comparable environment would justify authoring. The result is limited to the model and harness information exposed to this manager, the active instruction and skill-discovery surface, the available tools, and the zero-sample preflight on 2026-07-18. It is neither universal evidence nor statistical proof, and no reviewer independence claim is made.