agentic-praxis-grimoire

APG4 Bootstrap v0.1 Evaluation

Outcome

APG4 authored and provisionally validated six canonical APG v0.1 skill leaves. Structural checks, deliberate scenario application, and independent review support the recorded procedures within their sampled limits. They do not establish harness discovery, clean causal improvement, statistical reliability, universal trigger precision, production readiness, or stable maturity.

Subsequent integration correction

External review returned correction-required because APG4 committed the canonical leaves under skills/ without adding the .agents/skills/ repository projection used by Codex discovery. APG4’s six procedures, 18 scenario dispositions, two corrections, and independent reviews remain accepted.

APG4A later supplied six relative symbolic links from .agents/skills/ to the canonical leaves. APG4 did not test repository discovery, explicit invocation, automatic invocation, or dogfooding through that projection.

Environment and Superpowers boundary

Superpowers remained installed globally throughout APG4 because other repositories still rely on it. APG repository instructions made it reference evidence only for this work and rejected its workflow authority unless a human- authorized task explicitly named a specific skill for inspection or comparison.

That policy is behavioral. APG4 does not claim that the plugin was unloaded, absent from model context, or mechanically disabled. No global configuration, alternate Codex home, plugin adapter, or isolated profile was created.

Method

APG4 did not retry the blocked APG3 A/B design. Instead, each skill received a frozen public-safe positive-use scenario, non-trigger scenario, and edge or stop scenario. One fresh read-only worker per skill deliberately read and applied the procedure, returned an artifact or disposition through the agent harness, and checked for ambiguity, ceremony, missing constraints, authority drift, and private-source leakage.

This was deliberate skill application, not automatic-trigger telemetry. No baseline condition was run, and no relative advantage over ordinary Codex or Superpowers was scored.

After scenario application, fresh reviewers that did not author the skills assessed trigger precision, authority, proportionality, completeness, project- policy separation, public/private safety, source-expression independence, evidence expectations, rollback clarity, and consistency with ADR 0003.

Results by skill

Skill Positive use Non-trigger Edge or stop Corrections Independent disposition
composing-bounded-worker-assignments Produced one bounded read-only audit assignment Declined a direct typo correction Stopped on unresolved authority and overlapping ownership 0 Accept provisional
designing-significant-changes Produced a reversible configuration-design recommendation with owner decisions preserved Declined an already accepted design Stopped before unauthorized permanent-deletion design 0 Accept provisional
planning-repository-work Produced four dependent implementation units and an integrated gate Declined one local link correction Stopped on unresolved key-custody design 1 Accept provisional after corrected-skill review
implementing-with-test-discipline Produced a defect-reproduction, narrow-fix, and fresh-verification sequence Declined behavioral ceremony for a README typo Stopped speculative production-only mutation 0 Accept provisional
debugging-systematically Produced a two-hypothesis bounded investigation and verification path Routed an established cause to implementation Stopped on secret-bearing evidence and unauthorized live mutation 1 Accept provisional after corrected-skill review
reviewing-and-verifying-repository-work Deferred acceptance until the asserted staged evidence was actually supplied Declined design brainstorming Stopped on stale tests, missing privacy evidence, and ambiguous dirty state 0 Accept provisional

Corrections

Repository planning

The initial stop language could have rejected safe sequential work merely because units touched overlapping files. The correction permits sequential units when overall ownership and integration are clear, restricts parallel work to independent scopes, and stops only when overall write authority or integration ownership is unresolved. All three planning scenarios were rerun without regression, and an independent reviewer accepted the corrected leaf.

Systematic debugging

The initial text called a correction “authorized by the evidence.” Evidence can support a correction but cannot grant mutation authority. The correction now requires both evidentiary support and permission from the current assignment and project-owned mutation authority. All three debugging scenarios were rerun without regression, and the original independent reviewer accepted the corrected leaf.

No skill required a second correction. No skill was omitted.

Independent review

All six skills passed the ten required review dimensions after the two bounded corrections. Review found no blocker, no remaining material skill defect, no delegation authorization, no human-scope expansion, and no managed-report requirement for ordinary internal workers.

Two non-blocking clarity notes remain for dogfooding: the implementation and review skills leave rollback or recovery requirements to the controlling bootstrap and project policy rather than naming them in every local parameter list. The review-positive scenario also withheld its synthetic diff and test output, so it tested evidentiary restraint rather than substantive finding quality. A concrete public-safe review fixture is required before any maturity claim beyond provisional.

Structural and provenance result

Every retained canonical skill is a direct child of skills/, contains only one SKILL.md, uses valid name and trigger-focused description frontmatter, defines explicit non-trigger and stop behavior, and has no unused support directory.

The bundle uses APG-native synthesis informed by generalized maintainer-authored practices, RepoMap contribution evidence, the public Agent Skills specification, and Superpowers under the MIT License. No skill copies or adapts a Superpowers template, slogan, diagram, rationalization table, fixed workflow chain, or project-specific command. APG’s own distribution license remains deferred.

Limitations

Next evidence

After APG4A, begin a new Codex session that can load the committed repository projection. Verify discovery before using the six skills deliberately where their triggers naturally occur in APG and in at least one additional real repository. Record positive and non-trigger decisions, artifacts or stop dispositions, corrections, authority and privacy findings, and unrun checks. Add a concrete public-safe review fixture and a schema or contract-change implementation case.

External review should accept, request bounded correction, defer, or reject the provisional bundle. It should not promote a skill, decommission Superpowers, or authorize another implementation phase through this record.