00003Inspect current project and source evidence, compare the credible first implementation choices under common criteria, correct the public authority model, preserve exact decision evidence outside the public projection, and propose exactly one bounded next implementation phase without implementing or authorizing it.
APG now states the complete authority chain:
Human maintainer
retains ultimate project, roadmap, publication, license, and
destructive-action authority
ChatGPT manager
plans and reviews only within a human-authorized task, phase, or
preapproved roadmap envelope
Top-level Codex manager
executes the bounded ChatGPT assignment, delegates internally, integrates,
validates, commits and pushes when authorized, and reports
Internal Codex workers
perform bounded assignments and return through the agent harness
ChatGPT may construct prompts and review reports inside the approved envelope, but it may not expand the roadmap, introduce an unapproved epic, approve destructive work, select publication or license terms, or grant itself broader authority. Each Codex layer may narrow but not expand its assignment. New epics, material scope expansion, publication, licensing, and destructive actions return to the human maintainer. Commits, worker results, reports, and recommendations remain evidence rather than automatic acceptance.
This correction is adopted under explicit APG2 authority and does not depend on the proposed roadmap sequence.
All candidates were assessed against demonstrated problem severity, dependency necessity, immediate APG value, observability, activation precision, mechanical stability, publication safety, provenance, maintenance, reversibility, Codex overlap, and sequence leverage.
Recommended as provisional APG3. A narrowly triggered skill can test the two-sided assignment problem: omitted boundaries and evidence versus duplicated policy and ceremony. It directly exercises APG’s manager-worker model and can produce matched baseline-versus-skill evidence. It must be allowed to reject itself when ordinary Codex is already adequate.
Deferred, not rejected. APG1 demonstrated real confidentiality and record risks, but the current public surface is small, clean, and still changing. Narrow rules are mechanical; identity, source topology, licensing, provenance, and authorization truth still require judgment. APG3 retains a mandatory manual gate. Reconsider automation before publication, after any repeated defect, when skill work creates repeated stable need, or at the next roadmap re-planning.
Deferred as a standalone foundation, not rejected. APG3 needs a minimum evaluation contract, but a reusable skill, fixture tool, or framework would be circular before APG evaluates one real skill. The contract is embedded in the vertical slice and reusable machinery is reconsidered after repeated need across later skill evaluations.
No fourth implementation prerequisite was identified. The authority correction is directly authorized documentation work inside APG2.
Provisional APG3 should evaluate whether to implement one direct-child skill
named composing-bounded-worker-assignments. It authors the leaf only after
calibration demonstrates a material baseline deficit. The skill would produce
exactly one proportional internal worker assignment only after delegation is
already permitted and independently selected by a manager.
The output contract always carries one objective, source scope, write scope or read-only status, material prohibitions, required evidence, concrete deliverable, observable acceptance, and the normal agent-harness return. Compact assignments may combine those semantic fields without eight headings. Snapshots, invariants, validation, commit authority, privacy, stop conditions, cleanup, and rollback appear only when an observable task condition makes them material.
The skill does not authorize or dispatch delegation, replace repository instructions, reopen an approved design, create external prompts, schedule workers, implement a manager runtime, or require managed internal-worker reports.
An independent evaluator must freeze calibration cases, sealed confirmation cases, bounded sampling, a rubric, a baseline-adequacy rule, reviewer contracts, and a safe execution subset before authoring. The rule names the qualifying binary failures and anchored-score deficits. The plan also freezes a finite behavior-family list, an overall sample maximum, and at most one reserve case for each affected family within that maximum. No-skill calibration must meet the frozen material-deficit rule before a leaf is authored. The candidate author may use calibration evidence and at most one correction, but cannot inspect sealed confirmation or reserve inputs before the candidate is frozen.
Baseline and candidate cases retain the same model, harness, repository instructions, overlapping skills, inputs, and tools; candidate availability or deliberate loading is the sole intended treatment difference. Activation cases make the skill available without forcing it. Composition cases state that delegation is already authorized and selected. The set covers eligible small, medium, higher-risk, and parallel assignments; trivial direct work; coupled work; ambiguous authority; and partial or blocked outcomes.
Evaluation requires correct positive and negative activation, no invented authority or forced delegation, precise non-overlapping ownership, complete semantic-core coverage without fabricated constraints, proportional length, correct harness return, and blind manager review of a substantive advantage over baseline. Downstream compliance is scored only from preselected matched fresh-worker runs in public-safe disposable or read-only fixtures; reviewer prediction is not compliance evidence. The frozen plan defines two independent confirmation reviewers, concealed condition labels, randomized pairing, anchored scoring, manager-acceptability labels, bounded reserve expansion, and disagreement and variance handling. Ties, confirmation failure, no clear advantage, or unresolved disagreement or variance require rejection or deferral. Frontmatter, leaf shape, links, progressive disclosure, provenance, and public safety receive focused checks.
If baseline calibration does not meet the frozen material-deficit rule, APG3 authors no leaf and records a truthful rejection or deferral. If a later candidate over-triggers, expands authority, leaks private conventions, requires managed worker reports, or adds ceremony without material benefit, APG3 removes or narrows the leaf and records a truthful rejection or revision. APG3 creates no dependent runtime, registry, taxonomy, adapter, or compatibility layer.
APG3 permits at most one bounded correction from calibration evidence followed by a complete calibration re-run. Confirmation failure, a second correction need, new problem class, or required rubric change ends the phase as rejected, deferred, blocked, or stopped.
APG2 used four bounded read-only assignments:
| Worker | Assignment | Result | Top-level disposition |
|---|---|---|---|
| A | Authority and protocol audit | Found material human/ChatGPT conflation and missing envelope stops | Accepted; public owners corrected |
| B | Manager-oriented skill analysis | Recommended a conditional, evaluation-first bounded-assignment skill | Accepted with proportional evaluation thresholds |
| C | Publication-validator analysis | Found manual gate sufficient for APG3 with explicit reconsideration triggers | Accepted; validator deferred |
| D | Skill-evaluation sequencing | Recommended an embedded minimum contract and deferred framework work | Accepted |
Workers returned through the agent harness. They created no managed reports, commits, pushes, or repository files. The top-level manager integrated only useful evidence and independently reviewed the source state.
Public documents use generalized source families, the canonical RepoMap identity, public Superpowers license facts, and APG-native synthesis. They do not expose internal repository identities, source topology, private commits, local paths, managed report destinations, or exact non-public scenario mappings.
Exact source snapshots, reorganized paths, worker finding summaries, candidate analysis, uncertainties, and provenance mappings are retained in tracked publication-excluded APG2 decision evidence. No public file links to that area. No raw worker output, chain-of-thought, credential, raw managed report, or private runtime state was preserved.
ADR 0002 remains Proposed. The roadmap records APG0 and APG1 as complete, APG2 as a completed proposal phase, and one proposed APG3 pending external disposition. Later topics remain unnumbered candidate themes.
The APG2 commit does not authorize APG3. No implementation begins while ADR 0002 remains Proposed.
The independent synthesis critic found no blocker and confirmed that exactly one APG3 is recommended, ADR 0002 remains Proposed, APG3 remains unauthorized, no implementation occurred, and no public/private leakage or copied external expression was evident.
It required three material corrections:
All three findings were accepted and corrected. The critic’s minor finding that sequence leverage was stated as an established ranking was also corrected to a provisional manager judgment.
A separate fresh read-only worker then reviewed the complete staged APG2 diff and returned accept with no blocker, material finding, minor finding, or requested correction. It confirmed that human and ChatGPT authority are distinct and bounded; exactly one evidence-supported APG3 is proposed; the alternatives are fair; the evaluation is observable and limited to one candidate, one variance expansion, and one correction cycle; APG3 contains no hidden second project; no public leak or dependency on publication-excluded evidence is evident; ADR 0002 remains Proposed; APG3 remains unauthorized; no implementation occurred; both counters are correct; and this exit truthfully requests external disposition.
A later completion audit found the APG2 result already committed on the private development line. It preserved that history and repeated all four read-only research tracks before preparing a forward-only documentation correction.
The authority audit found one material ambiguity in the root precedence rule: delegated direction could appear to outrank repository instructions even though the normative protocol reserves supersession to explicit human-maintainer authority. The skill and evaluation audits found three material contract gaps: the assignment core appeared optional, invocation was not observable, and downstream compliance was scored without actual execution. They also found that one visible scenario set allowed correction overfit and that a fixed sampling matrix added ceremony without supporting a general statistical claim.
The corrections name human-maintainer precedence, retain all eight semantic assignment fields without mandating eight headings, separate unforced activation from deliberately loaded composition, require bounded disposable execution for downstream claims, and separate calibration from sealed confirmation with one pre-frozen reserve expansion bounded by per-family and overall caps. Sampling is predeclared and bounded, final claims remain sample-specific, and the one- correction limit remains intact.
The publication-validator audit found no current leak, broken public local link, or record defect. It confirmed that the complete manual gate remains proportionate through APG3 while recent content churn and judgment-heavy classification make immediate validator code premature. The validator remains deferred, not rejected.
The fresh synthesis critic found two further material gaps and one minor ambiguity. The no-leaf fork lacked a predeclared material-deficit rule, matched results were required even when no leaf was authored, and reserve scope was not internally precise. The corrections freeze a baseline-adequacy rule, make matched confirmation conditional on leaf authoring, and bind a finite behavior- family list and per-family reserves to one overall sample maximum. The critic found no authority expansion, hidden framework, private leak, acceptance bypass, implementation, or authorization drift.
The staged documentation-only gate produced these results:
| Check | Result |
|---|---|
| Development branch and remote baseline | Passed; clean fetched main began equal to origin/main, and all 20 staged APG2 paths are Markdown. |
| Read-only source repository | Passed; main remains clean, unchanged, and equal to its fetched remote snapshot. |
| Complete phase scope | Passed; changes are limited to authority/governance documents, ADR and exit records, roadmap/skill/provenance indexes, and publication-excluded decision evidence. |
| Markdown structure, fences, and local links | Passed across 29 tracked Markdown files and 64 links; no broken local target or odd fence count. |
| ADR namespace | Passed; unique independent sequence 0001, 0002, with index coverage. |
| Exit namespace | Passed; unique independent sequence 00001, 00002, 00003; every filename ends in -exit.md, every title contains Exit, and index coverage is complete. |
| Public confidentiality scan | Passed across 18 projected paths; no private repository identity, source topology, private snapshot, private commit, or user-specific absolute path. The only full hash is the classified public Agent Skills pin. |
| Report-destination classification | Passed; only generic executable-interface paths and policy descriptions remain public, with no private managed report destination. |
| Public/private dependency check | Passed; no publishable Markdown link enters the publication-excluded tree. |
| Authority hierarchy | Passed; human ultimate authority, delegated ChatGPT scope, Codex narrowing rules, repository-decision precedence, and reserved-decision stops are explicit. |
| Proposed status and authorization | Passed; ADR 0002 is Proposed, the roadmap contains exactly one proposed APG3 heading, and APG3 is explicitly not authorized, active, or accepted. |
| No implementation | Passed; no SKILL.md, executable validator, framework, source/test code, dependency metadata, runtime, scheduler, registry, package, or publication artifact was added. |
git diff --check |
Passed. |
git diff --cached --check |
Passed. |
| Source tests and compilation | Not run; docs-only. |
The complete staged diff passed the separately mandated fresh finished-diff review with no finding.
The forward-only documentation correction received a separate final gate. The original 20-path gate above is historical evidence and is not evidence for this later diff.
| Check | Result |
|---|---|
| Development and reference baselines | Passed; development began clean and equal to its fetched remote, and the read-only reference remained clean, unchanged, and equal to its fetched remote. |
| Forward-correction scope | Passed; all six changed paths are Markdown and are limited to root authority wording, the proposed ADR, roadmap and skill indexes, the APG2 exit, and publication-excluded dispositions. |
| Markdown structure, fences, and local links | Passed across 29 tracked Markdown files, 64 links, and 62 local links; no H1 count, heading-level, fence, target, or public-to-private link error. |
| Public confidentiality scan | Passed across 18 projected paths; the only two full-hash occurrences are the classified public Agent Skills pin, with no private commit identity, managed report identifier, user-specific absolute path, or local username. |
| ADR and exit namespaces | Passed; ADR sequence 0001, 0002 and exit sequence 00001, 00002, 00003 remain unique and complete. |
| Proposed status and authorization | Passed; ADR 0002 is Proposed, the roadmap contains exactly one proposed APG3 heading, and APG3 remains unauthorized. |
| No implementation | Passed; no skill leaf, executable validator, evaluation framework, source or test code, dependency metadata, runtime, or publication artifact was added. |
git diff --check |
Passed. |
git diff --cached --check |
Passed. |
| Source tests and compilation | Not run; docs-only. |
APG2 changed documentation and tracked publication-excluded decision evidence only. It implemented no skill, validator, framework, runtime, scheduler, registry, dependency, source code, test code, publication automation, public repository, or release.
APG3 is not authorized.
ChatGPT should review the APG2 final response, exact commit report, operational report, ADR 0002, roadmap, and this exit within its delegated authority. The human maintainer must supply the decision when no existing roadmap envelope already delegates it.
The requested disposition is exactly one of:
A later bounded acceptance or correction phase must record that disposition and change ADR 0002 from Proposed before any APG3 implementation begins.
APG2A later accepted ADR 0002 and authorized the bounded APG3 vertical slice. The APG2A acceptance exit records that later disposition. This section does not rewrite the historical APG2 close: ADR 0002 was Proposed, APG3 was unauthorized, and no implementation had occurred when APG2 completed.