Staged agent workflows
Centralize control, not raw context or execution.
Session lifecycleβ
Scope β Understand β Produce β Verify β Deliver
The active root owns this sequence. Skills contribute task mechanics and acceptance criteria; they do not add nested workflows. Empty stages need no ceremony. Verification occurs once on the integrated, formatted candidate. SOP Β§3.5 lets the root diagnose and repair failed checks within existing authority, freeze the repaired candidate, and rerun failed and affected checks. SOP Β§3.4 stops repeated attempts without new evidence or progress; workers never own recovery loops. Convergence is explicit-only and requires a finite user-approved repair/check allowance before entry.
Context and model responsibilitiesβ
| Category | Responsibility | Context |
|---|---|---|
| orchestrate | Strong root: intent, decisions, dependencies, integration, stage transitions | Compact task handoff, not all source/logs |
| research | Strong bounded investigation | Task-specific source; return conclusions, evidence pointers, uncertainty, affected interfaces |
| implement | Cheaper implementation band: substantial settled edits | Owned targets and ready design inputs; return unverified artifacts |
| mechanical | Deterministic tools directly; cheap model only when needed | Exact rule and targets |
| review / refute | Strong final judgment, selected framing and independent risk lenses | Actual candidate source plus shared check receipts |
| memory | Automatic staged recall admission and final batched learning | Compact admitted lines; no per-turn memory agents |
home/.chezmoidata/ai_models/tiering.yaml remains the model/effort authority. This change preserves model selections. Substantial routine implementation must not default to the expensive root/review model. A user-requested inline session is the explicit exception.
Worker interfaceβ
A packet names stage/category, question/change, owned targets, ready evidence, applicable project/safety constraints, named role mechanics, intended/preserved differences, output, forbidden effects, and terminal condition. Pass the needed constraints explicitly, not the whole SOP, instruction tree, catalog, or parent transcript. Missing required constraints block the packet. Independent work may run concurrently; dependent work starts only when inputs exist. Workers never spawn models, message siblings, broaden scope, run private QA, or reopen after completion. Return produced or blocked with artifact pointers; production does not return green/approved verdicts. Final specialists return findings/evidence once and never verify one another. Keep the substantive terminal result immutable. Late events cannot replace it with status chatter.
Long sessionsβ
Keep raw source, diffs, search output, and logs outside root context. Persist a compact handoff in the existing active topic: stage, scope/snapshot, settled decisions, dependencies, active/completed packet IDs, open questions, and evidence pointers. After compaction, continue from it without rediscovery or relaunch. Final reviewers still inspect actual relevant source, not summaries alone. Known deterministic commands use tools directly; a separate agent per read/check wastes context without adding judgment.
Preserved responsibilities, new ownersβ
Removing child orchestration must not remove the work it used to carry. The migration from the pre-stage contracts (bc9da0992f5f) keeps these responsibilities at the following call sites. These are stage owners, not a mandatory agent roster.
| Previous responsibility | Current owner and reachable contract | Deliberately removed machinery |
|---|---|---|
| Necessity, intent, forks and acceptance packet | Root Understand: k-spec β check-strength / packet-template; verified overlays supply domain planning forks | Automatic necessity/advisor/prototype ceremonies and red-check approval loops |
| Large code/public-source investigation and design alternatives | Strong Understand packets: code-searcher, public-sources, k-codebase-design β going-deeper; root owns decisions | Per-query/per-alternative orchestration; cheap research substitution |
| Settled implementation, tests, docs and generation | Root Produce β implementation-band implement-worker; direct tools for deterministic transformations | Worker self-review, test-to-green and mechanical check-runner agents |
| Criterion truth, user-path reachability, clean-state durability and scope accounting | Root final Verify: k-build β criteria-verifier over actual source and shared receipts | A verifier per criterion, receipt re-execution and mandatory mutation of every branch |
| Low-risk eligibility and complete judging rules | Root routing: k-light-review predicate β judging_core + judging_pipeline | Treating a small diff as low risk; light finder/auditor chains |
| Expert criteria, complete assigned coverage and incidental real defects | Root k-review / deep reviewer-roster β lanes; final reviewer-worker receives selected checks and reports coverage/gaps | Agent per heading; stopping at the first severe finding; leaves loading the roster |
| Artifact review and adversarial challenge | Distinct strong final questions against the same candidate; registry review/refute bands and adversarial-verifier | Reviewing another review as independence; weaker-family substitution |
| Redundancy, verbosity, semantic duplication, missing consumers/docs/tests | Integrated judging_pipeline, reached by standard/deep/light judging | Separate post-review/hygiene certification pass |
| Finding conflicts, material missing evidence and blind clarity | Final synthesis in judging_pipeline; conditional fresh-eyes retains its blind packet | Model votes, silent blocker deletion, or PR narrative used to dismiss newcomer confusion |
| Public claims and exact numeric/source support | Root k-public-sources β final batched claim-verifier; complete captured sources, exact quotes and URLs | Per-claim verification and collectβverifyβdeepen cycles |
| PR necessity, current intent and correctly-open status | Root Understand: deep pr-necessity / standard pr_common β pr_context_audits | Necessity controller ladder; spending on superseded work without an explicit reason |
| Live UI applicability, branch/config/data truth and screenshots | Final live-ui-validation β live-ui-review β live-ui-runtime; verified overlays own target/setup policy; root views used images once | Source fixes inside UI verification and automatic post-judgment fix tasks |
| Snapshot, discussion, pending-review deduplication and publication | Root pr_snapshot / pr_common / review_delivery and publication skill | Worker PR refetches; review authorship treated as edit or publish authority |
| Diagnosis, cleanup and PR-fix batches | Root authorizes known production through k-diagnosing-bugs, k-knip, k-pr-fix-loop / pr_fix; SOP owns integrated final checks and scoped recovery | Private check loops, unrequested class-wide cleanup and automatic new-comment drains |
| Recall admission, corrections and durable learning | Root k-ai-kb β smol-operator / cli; staged recall and one final learning batch; identical inline fallback | Per-turn scribes, leaf memory orchestration and dropping learning when delegation is forbidden |
| Convergence and prose alternatives | Explicit k-converge invocation (declared exit + correctness filter) or requested k-text-tournament; root owns the enclosing lifecycle | Automatic convergence and repeated evaluator tournaments |
Existing named profiles in Claude, Codex, Cursor, Copilot, Pi and OMP point to the same leaf contracts. Dynamic or unsupported adapters remain subject to the capability boundaries below; a contract reference does not prove profile discovery or runtime enforcement. Contract tests check required load edges, profile bindings and retained obligations. Inline source/evidence judgment remains necessary: text tests do not prove model compliance, lower token use or large-project quality parity.
Runtime controls and limitsβ
Coverage inventoryβ
Coverage is the join of the launcher, managed configs, provider wrappers and installed application surfaces. ,ai is not the harness inventory: it routes seven native CLIs and omits OMP and Crush. A native CLI check does not certify a subscription wrapper, local provider, hosted frontend or desktop application.
| Surface | Configured scope and confidence boundary |
|---|---|
| Claude Code, Codex, Cursor CLI, Copilot CLI | Native profiles/hooks plus the separately listed provider routes. Native child controls and model selection are distinct claims. |
| Pi, OMP | Native packages/extensions and role profiles; see controls below. |
| OpenCode | Existing work implementation workers; personal single-context. Strong category lanes are not configured. |
Antigravity CLI (agy, registry key gemini) | Native dynamic tiers and shared hooks. This is not evidence for a separate Gemini CLI or the desktop application. |
| Crush | Managed crushrc loads the SOP. Installed v0.92.0 task defaults are read-only and share the large model; no managed category lanes or automatic memory hooks. Root-only native hooks do not certify child lifecycle enforcement. |
| OpenRouter wrappers | Claude, Codex, Copilot and Cursor wrappers select the Pi backend matrix. Claude and Codex project managed native profiles with exact selectors, including distinct refute effort; the gate denies missing or stale role maps. |
| Codex subscription wrappers | ,claude-codex, ,copilot-codex, ,cursor-codex; use the Codex backend matrix, not the frontend's model catalog. |
| Copilot subscription wrappers | ,claude-copilot, ,codex-copilot, ,cursor-copilot; use the Copilot backend matrix and entitlement-filtered catalog. |
| llama.cpp wrappers | Claude, Codex, Cursor and OpenCode launchers. A single local model is not a strong/cheap category ladder; native lifecycle controls still belong to the frontend. |
,q | Deliberately stripped Pi one-shot: no discovered extensions, skills, context or session persistence. It is not the staged coding workflow and does not inherit the Pi extension guards. |
| Cursor / Antigravity desktop | Separately installed application surfaces; CLI evidence does not establish desktop behavior. |
| tuicr / lgtm | Review interfaces, not additional model orchestration harnesses. |
Subscription adapters select child controls only through explicit @lane-<effort> selectors; a raw model ID never implies a worker role, even when it matches another registered lane. Raw root requests retain launch controls. Claude projects managed --agents definitions with exact registered pairs and preserves the profile body and skills. Selectors are stripped before the upstream request. The band hook removes call-level model overrides and denies missing/unavailable profiles, stale role maps, resume/fork requests, or conflicting controls. Projected tool lists exclude orchestration tools; root terminal-wakeup prevention remains a separate criterion.
,codex-copilot projects entitled lane selectors into a session-local Codex catalog and copies managed leaf profiles without native model/effort overrides. Codex applies profiles after spawn arguments, so both projections are required. The gate admits only projected roles and exact lanes; it rejects full-history forks, missing session projections, conflicting provider controls, and child-originated delegation. The profile copies retain their instructions and disable multi_agent. The session-local files live until the wrapper exits; native profiles remain unchanged. Relaunch sessions started before this projection existed.
The review runtime caveats in ~/.agents/skills/k-review/references/runtime-harnesses.md use this conditional support, not a blanket ,codex-copilot ban. Native child-tag transport and live provider completion are separate acceptance conditions: proving the former does not prove that a live refuter will complete.
Native child-tag transport is verified absent on ,copilot-codex: Copilot CLI 1.0.83 takes the child model from the per-subagent ~/.copilot/settings.json entry after the band hook rewrites the call's model, so an injected @lane-<effort> selector never reaches the loopback request (probed 2026-09-11).
Delegation itself now runs on that route, unpinned, because the adapter restores the turn's output items on the streamed terminal event.
,cursor-codex and ,cursor-copilot remain unverified.
Root sessions and native harness routes remain available. Fake-provider translation checks alone do not establish frontend selector acceptance.
Native Claude alias overrides are clamped in both directions: a cheaper alias must not weaken a strong role. Generic Codex/Copilot workers preserve a registered model/effort pair, not model membership alone. A model shared by multiple lane efforts requires an explicit matching effort. Cursor uses a model selector instead: its registered auto mechanical/memory rows intentionally omit effort. Backend-schema routes validate the backend contract, not the frontend's exception. Missing or malformed required lane data denies delegation on the configured mutating adapters. Copilot propagates those denials and rejects delegation when its band helper is missing, fails, or emits malformed output; optional memory errors remain independent.
Native controlsβ
- All profile templates include the shared leaf contract. Former controller profiles are final judgment leaves, not nested controllers. Pi/OMP reviewer and blind-clarity profiles name the active root as their dispatcher; reviewers must not reload an ambient catalog or treat a controller-named sibling as an orchestrator.
- Pi profiles set
defaultContext: fresh,inheritProjectContext: false,inheritGlobalContext: false,inheritSkills: false, and child depth zero. This removes ambient instructions/catalog, not explicitly named role skills: pi-subagents 0.66.0 loads those separately inruns/foreground/execution.ts. The root supplies applicable project and safety constraints in the packet. Dispatch requiresagentScope:"user"so project profiles cannot override managed models, tools, or inheritance. The runtime guard requiresacceptance:falseand rejects context/model/per-call skill overrides, composite workflows, host gates, scheduling and resume/steer recovery. Dispatch one ready packet per call; independent root packets may explicitly setasync:true. Defaults are fresh and foreground. Native child identity suppresses root memory/reinforcement injection and blocks the subagent tool. - Claude generic implementation profiles omit agent tools. Unrestricted shell remains a limitation, not a sandbox guarantee.
- OMP retains native depth pruning and disables background advisors. Explicit root async and effort controls remain available; implicit Bash/Eval backgrounding is disabled because it also affects blocking workers. Managed profiles (including native-name
task,sonic, andscoutshims) useblocking: true, notasktool, and nospawnsallowlist; independent task batches still run concurrently before returning. Native reviewer shortcuts are disabled in favor of managed review profiles. - OMP's runtime extension latches the native
yieldcapability as the leaf boundary and rejects leaf async, task, advisor and hub calls. This prevents those tool paths from creating work that outlives the return. Explicit-yield root sessions receive the same conservative restrictions. Ordinary roots retain named-process hub operations; peer sends remain blocked with native whitespace normalization. The native driver is not patched: an externally created pending job can still invalidate a yield. Do not confuseblocking:truewith disabling background jobs or claim universal terminal enforcement. - OMP 18.1.14 passes context files, skills, and native child/Coop instructions through
src/task/executor.ts; its profile parser does not expose Pi's inheritance flags. The managedbefore_agent_starthook replaces the leaf model prompt with its role, explicit packet/context/plan, native worktree restriction, and yield protocol/schema. It removes the inherited SOP/catalog and native private-QA/peer-wakeup instructions; named skill autoload messages remain. Unknown, ambiguous, or unmanaged prompt frames abort before provider dispatch instead of falling back to a controller prompt. This controls the model-facing prompt, not physical SDK discovery, native revival, or later third-party prompt rewrites. Do not use an adapter unattended when it cannot enforce the required no-orchestration/terminal boundary. - Former controller leaves omit declared edit/write/agent/task tools where supported. Cursor retains
readonly: falsefor its existing shell/MCP access caveat; prompt-level read-only instructions are not runtime write isolation. - Shared startup/per-turn hooks skip Codex/Claude child payloads carrying
agent_idbefore topic lookup or root-context injection. Root hooks retain filtered KB retrieval and staging; only the root owns admission and final learning. No per-turn scribe or automatic convergence. Topic binding, worklogs, context-disable sentinels, and reinforcement remain. - Codex disables
features.multi_agentin every managed child role, including nativedefault,worker, andexploreroverrides. Generic roles leave model selection to the existing band hook, preserving explicit registry-backed lane picks; dedicated roles retain their declared bands. Root collaboration remains enabled. This is CLI role configuration, not proof that a hosted frontend honors local role files. - OpenCode work retains its existing implementation workers/models, now with the shared leaf prompt and native
permission.task: deny. Only the work root may dispatchworker_*; personal remains single-context. The newly available built-in scout is disabled alongside existing built-ins. The memory plugin resolves nativeparentID: root recall/learning plumbing remains, children keep worklogs without root recall, and child/unknown-identity task calls fail closed. Existing worker-session resume requests are rejected. OpenCode still has no category-backed strong research/review/refutation adapter; do not substitute its implementation workers for those stages. - Cursor/Copilot retain their model-band adapters. Claude/Copilot managed profile tool allowlists omit delegation. Cursor CLI does not discover user custom profiles, and Antigravity's native dynamic tiers are not leaf lifecycle enforcement. These surfaces still require the root's packet/no-resume discipline; do not run unattended isolated work where the runtime boundary is unavailable. Unrestricted shell/MCP is not an agent sandbox in any harness.
OMP 18.1.14 dispatches blocking profiles through its synchronous fan-out path (src/task/index.ts); src/discovery/helpers.ts parses the flags. The root guard uses native pi.pi.discoverAgents(cwd) before dispatch. Task packets must name their profiles explicitly, including every batch item; implicit native defaults are rejected. Each selected definition must resolve to its enabled user-profile file under the canonical agent directory and carry blocking, explicit model/tool and leaf-boundary controls. Project, plugin-path and bundled replacements are rejected. Settings-level model overrides, worker advisors and prewalk handoffs are rejected too; explicit off values and final refute profiles using @advisor remain available. Eval can select profiles dynamically, so its admission checks every enabled definition. This can block non-agent Eval code when an enabled definition is unsafe; use ordinary root tools or correct the profile, not an unguarded dispatcher.
OMP CLI task/Eval dispatch also checks effective Bash/Eval foreground settings immediately before dispatch. Project/CLI overlays can override the managed file; unsafe or unavailable values block dispatch with the offending keys. The guard reads the CLI's bundled pi.pi.Settings.instance and reloads disk layers through the same native method as subagent preflight; it sets no values or overrides. A concurrent terminal-marker failure also blocks an admission awaiting discovery. Source/legacy-module imports can resolve a different singleton and silently return defaults in the shipped npm CLI. This is a CLI-root control, not certification of secondary roots created with isolated SDK settings. Those roots remain unsupported for unattended dispatch.
Eval has a separate native bypass. In OMP 18.1.14, eval/js/tool-bridge.ts routes __agent__, __workpool__ and __completion__ before ordinary tool lookup and wrapping; both Eval runtimes use that bridge. agent() registers background work with keep-alive execution in eval/agent-bridge.ts, irrespective of a profile's blocking flag. workpool() creates persistent pools; completion() resolves its own tier/effort and starts a model request without a managed profile. Passing the outer Eval admission guard does not enforce the inner helper's lifecycle or model lane. These helpers and their synthetic bridges are prohibited as managed dispatch routes; use explicit native task packets instead. Ordinary Eval is not disabled, and no runtime interception of these helper calls is claimed.
The k-omp adapter now puts task dispatch and named-process hub operations under its root-only section. A leaf skips that section rather than receiving a generic instruction to perform agent handoffs. The root-moves instruction scanner recognizes the quoted native task command as well as βTask toolβ; its counterexamples retain prohibitions, ordinary task-state prose and root-gated dispatch.
Native OMP plan-mode dispatch is a confirmed unsupported path: structured-subagent.ts sets restrictToolNames, and sdk.ts loads an empty extension list for those children. That skips the managed leaf prompt/tool/terminal hooks. A per-call Eval-defined tools list alone does not set this restriction. The public extension context has no plan-mode getter, while ACP and plan-yolo can change mode without the TUI's transcript entries; transcript inspection would not close the gap. Do not dispatch unattended workers in native plan mode or restricted SDK sessions. This restriction does not change the strong @plan model role or prohibit root planning. Profile admission is not a fix for extension omission, concurrent external file edits, or altered files already inside the user profile directory.
A native yield while owned asynchronous jobs remain is provisional: the executor may collect results and request another yield before returning to the parent. A test observing two yields before the parent receives a result proves additional internal turns, not post-completion resurrection. Terminal tests must observe the parent-visible result and subsequent delivery separately; neither a one-yield fixture nor the foreground-settings guard proves every native revival path is sealed.
The OMP root also consumes the native task:subagent:lifecycle finalized event. Only completed, failed, or aborted creates an exclusive <sessionFile>.k-leaf-terminal marker beside that worker's native transcript. The native finalizer emits this after saving and classifying the result, not at a provisional yield. The marker leaves the transcript and result untouched. Managed leaves check it before prompts, provider requests and tool execution, including a cold reopen of the same session. This prevents guarded execution after the observed final event; it does not prevent native registry retention or session construction. Leaves without persisted session identity are blocked. A failed marker write makes the root task/Eval receipt an error and blocks further dispatch; other completed workers are still sealed.
Do not treat pre-existing workers without markers, removed/relocated markers, disabled/replaced extensions, or arbitrary shell/MCP as covered by that boundary. Do not resume a completed workpool batch as a new packet. The native-pipeline fixture uses a synthetic leaf identity and finalized event, so full task/revival confidence still requires a separately authorized native worker pilot.
OpenCode's task/no-resume guards register even when optional memory helper files are absent. Each available memory callback remains active independently; missing recall or worklog plumbing does not turn off lifecycle controls.
Prompt contracts do not prove native enforcement. Unsupported autonomous lifecycle controls require a visible capability limitation rather than a claimed guarantee. No universal spend cap is asserted; usage accounting must include children/advisors when the harness exposes it. The policy evaluation matrix includes every capability-snapshot harness, including OMP and Crush; runtime-only child bindings are reported unsupported for static-pinning scenarios, not silently omitted.
These controls prevent specific sources of extra work; they do not classify arbitrary shell commands as research, production or verification. The root still owns the single final check plan. A command allowlist broad enough for implementation is not proof that a worker cannot run private QA. Runtime boundary tests and a small worker pilot also do not establish better quality or speed on a matched Kibana-scale workload.
Sourcesβ
Core: home/readonly_AGENTS.md Β§Β§3.5β3.7. Leaf: home/dot_config/exact_tmux/agent_prompts/leaf-boundary.txt. Recipes: home/exact_dot_agents/exact_skills/. Runtime: shared hooks plus Pi/OMP extensions and profile templates. See model tiering, spec/build, and cross-agent memory.
Native configuration references: Codex custom agents βloads these files as configuration layers for spawned sessionsβ; OpenCode task permissions βControl which subagents an agent can invoke via the Task tool with permission.task.β Installed source remains authoritative for the tested runtime version.
Crush source: v0.92.0 hook scope: βSub-agents (the agent task tool, agentic_fetch, etc.) run without hook wrapping on their internal tools.β