agentic-praxis-grimoire

APG3 Bounded Worker-Assignment Skill Evaluation

Outcome

APG3 stopped at its mandatory harness preflight. The current internal-agent harness could create fresh agents and return bounded results, but it did not expose a skill-invocation event or a supported per-agent skill profile. An overlapping delegation and assignment-composition skill collection was active, so baseline and candidate conditions could not represent the accepted target profile without global mutation, an adapter, or an invented substitute.

Accepted ADR 0002 requires APG3 to stop under those conditions. No evaluator was assigned, no evaluation material was created, no baseline was run, and no candidate was authored.

Objective and authority

APG3 was authorized to evaluate whether ordinary Codex assignment composition has a material deficit and, only if a frozen baseline-adequacy rule demonstrates that deficit, author and evaluate at most one reversible composing-bounded-worker-assignments candidate. The candidate could apply only after delegation was already authorized and independently selected.

The authorization allowed APG3 to narrow or stop for evidence or environment limitations. It did not authorize a harness adapter, evaluation framework, plugin mutation, second skill, validator, runtime, dependency, taxonomy, publication decision, or licensing decision.

Environment preflight

The bounded environment class was the current Codex desktop internal-agent harness with the repository instructions and installed skill-discovery surface available on 2026-07-18. The candidate leaf was absent. An overlapping external skill collection was enabled in the manager environment.

The preflight produced this capability disposition:

Required property Result
Fresh agent contexts Available, but they inherited the same skill and plugin surface.
Bounded agent-harness returns Available.
Observable skill-invocation event Unavailable on the exposed harness event and tool surface.
Candidate available without forced invocation No supported per-agent availability control was exposed.
Deliberate candidate loading No supported internal-agent loading control was exposed.
Identical baseline and candidate environments Not credibly maintainable without the missing controls.
Overlapping-skill isolation No supported isolated internal-agent profile was available.
Sealed-input and downstream execution stages Not reached after the mandatory stop.

No evaluation environment was frozen because the hard preflight failed. An available command-line profile flag did not supply the missing internal-agent profile, invocation event, or treatment controls, and APG3 was not authorized to construct an adapter or weaker proxy.

Authorized behavior families

The accepted contract required eligible families for small read-only work, isolated-write work, contract-sensitive work, disjoint selected worker scopes, and partial or needs-context outcomes. It also required non-trigger families for trivial direct work, unauthorized or unselected delegation, unresolved authority or design, coupled ownership, an already adequate assignment, and external prompts, scheduling, monitoring, managed reports, or manager-runtime requests.

These families remained authorization requirements only. No case text was created and no family was sampled because evaluation-role assignment occurs after the failed harness gate.

Evaluation method and baseline rule

The accepted method would have separated unforced activation from deliberate composition, frozen distinct calibration and sealed confirmation evidence, bounded all samples and reserves, and required actual matched worker execution for downstream-compliance claims. The authoring gate would have required either a recurrent named binary failure or qualifying deficits across at least two anchored quality dimensions, or a qualifying manager-acceptability failure across a frozen share of eligible evidence. A lone outlier, minor correction, verbosity difference, or heading count could not qualify.

That method was not instantiated. No plan, rubric, baseline-adequacy rule, reviewer contract, calibration bundle, sealed bundle, reserve bundle, sample minimum, sample maximum, or execution subset was frozen. The executed sample count and reserve use were both zero. The baseline rule was not applied, so this outcome is not baseline-adequate-no-skill.

Results

Stage Result
No-skill calibration Not run.
Authoring decision Not permitted because preflight did not establish a valid experiment.
Candidate calibration Not applicable.
Activation Not run; invocation was not observable.
Sealed confirmation Not applicable.
Confirmation reviewers Not assigned.
Reserve expansion Not used.
Downstream execution Not run; no downstream-compliance claim is made.
Manager acceptability Not scored.
Binary evaluation safety Not evaluated because no samples ran.

The repository-level stop was safe: no candidate directory, empty placeholder, evaluation framework, fixture framework, validator, runtime, registry, scheduler, adapter, dependency, taxonomy, publication artifact, or licensing decision was added.

Skill and rollback state

No candidate was created, so no candidate hash, correction, freeze, rejection snapshot, adoption, or rollback mutation exists. The active skill tree remains without a SKILL.md. The terminal skill disposition is not authored because APG3 was blocked before evaluation.

Re-evaluation condition and external decision

Credible re-evaluation requires a supported internal-agent profile that removes the overlapping bootstrap and assignment-composition skills in both conditions, exposes an observable skill-invocation event, controls unforced availability and deliberate loading, and otherwise preserves model, harness, instructions, tools, and skill discovery. Providing or selecting such an environment is an external roadmap decision; APG3 did not create it.

External review is requested to accept the blocked disposition or authorize a separate roadmap decision about a supported evaluation environment. No later implementation phase is authorized by this result.

Limitations

No claim is made about ordinary assignment adequacy, candidate quality, activation safety, downstream compliance, or whether a future comparable environment would justify authoring. The result is limited to the model and harness information exposed to this manager, the active instruction and skill-discovery surface, the available tools, and the zero-sample preflight on 2026-07-18. It is neither universal evidence nor statistical proof, and no reviewer independence claim is made.