Skip to main content

Staged agent workflows

Centralize control, not raw context or execution.

Session lifecycle​

Scope β†’ Understand β†’ Produce β†’ Verify β†’ Deliver

The active root owns this sequence. Skills contribute task mechanics and acceptance criteria; they do not add nested workflows. Empty stages need no ceremony. Verification occurs once on the integrated, formatted candidate. SOP Β§3.5 lets the root diagnose and repair failed checks within existing authority, freeze the repaired candidate, and rerun failed and affected checks. SOP Β§3.4 stops repeated attempts without new evidence or progress; workers never own recovery loops. Convergence is explicit-only and requires a finite user-approved repair/check allowance before entry.

Context and model responsibilities​

CategoryResponsibilityContext
orchestrateStrong root: intent, decisions, dependencies, integration, stage transitionsCompact task handoff, not all source/logs
researchStrong bounded investigationTask-specific source; return conclusions, evidence pointers, uncertainty, affected interfaces
implementCheaper implementation band: substantial settled editsOwned targets and ready design inputs; return unverified artifacts
mechanicalDeterministic tools directly; cheap model only when neededExact rule and targets
review / refuteStrong final judgment, selected framing and independent risk lensesActual candidate source plus shared check receipts
memoryAutomatic staged recall admission and final batched learningCompact admitted lines; no per-turn memory agents

home/.chezmoidata/ai_models/tiering.yaml remains the model/effort authority. This change preserves model selections. Substantial routine implementation must not default to the expensive root/review model. A user-requested inline session is the explicit exception.

Worker interface​

A packet names stage/category, question/change, owned targets, ready evidence, applicable project/safety constraints, named role mechanics, intended/preserved differences, output, forbidden effects, and terminal condition. Pass the needed constraints explicitly, not the whole SOP, instruction tree, catalog, or parent transcript. Missing required constraints block the packet. Independent work may run concurrently; dependent work starts only when inputs exist. Workers never spawn models, message siblings, broaden scope, run private QA, or reopen after completion. Return produced or blocked with artifact pointers; production does not return green/approved verdicts. Final specialists return findings/evidence once and never verify one another. Keep the substantive terminal result immutable. Late events cannot replace it with status chatter.

Long sessions​

Keep raw source, diffs, search output, and logs outside root context. Persist a compact handoff in the existing active topic: stage, scope/snapshot, settled decisions, dependencies, active/completed packet IDs, open questions, and evidence pointers. After compaction, continue from it without rediscovery or relaunch. Final reviewers still inspect actual relevant source, not summaries alone. Known deterministic commands use tools directly; a separate agent per read/check wastes context without adding judgment.

Preserved responsibilities, new owners​

Removing child orchestration must not remove the work it used to carry. The migration from the pre-stage contracts (bc9da0992f5f) keeps these responsibilities at the following call sites. These are stage owners, not a mandatory agent roster.

Previous responsibilityCurrent owner and reachable contractDeliberately removed machinery
Necessity, intent, forks and acceptance packetRoot Understand: k-spec β†’ check-strength / packet-template; verified overlays supply domain planning forksAutomatic necessity/advisor/prototype ceremonies and red-check approval loops
Large code/public-source investigation and design alternativesStrong Understand packets: code-searcher, public-sources, k-codebase-design β†’ going-deeper; root owns decisionsPer-query/per-alternative orchestration; cheap research substitution
Settled implementation, tests, docs and generationRoot Produce β†’ implementation-band implement-worker; direct tools for deterministic transformationsWorker self-review, test-to-green and mechanical check-runner agents
Criterion truth, user-path reachability, clean-state durability and scope accountingRoot final Verify: k-build β†’ criteria-verifier over actual source and shared receiptsA verifier per criterion, receipt re-execution and mandatory mutation of every branch
Low-risk eligibility and complete judging rulesRoot routing: k-light-review predicate β†’ judging_core + judging_pipelineTreating a small diff as low risk; light finder/auditor chains
Expert criteria, complete assigned coverage and incidental real defectsRoot k-review / deep reviewer-roster β†’ lanes; final reviewer-worker receives selected checks and reports coverage/gapsAgent per heading; stopping at the first severe finding; leaves loading the roster
Artifact review and adversarial challengeDistinct strong final questions against the same candidate; registry review/refute bands and adversarial-verifierReviewing another review as independence; weaker-family substitution
Redundancy, verbosity, semantic duplication, missing consumers/docs/testsIntegrated judging_pipeline, reached by standard/deep/light judgingSeparate post-review/hygiene certification pass
Finding conflicts, material missing evidence and blind clarityFinal synthesis in judging_pipeline; conditional fresh-eyes retains its blind packetModel votes, silent blocker deletion, or PR narrative used to dismiss newcomer confusion
Public claims and exact numeric/source supportRoot k-public-sources → final batched claim-verifier; complete captured sources, exact quotes and URLsPer-claim verification and collect→verify→deepen cycles
PR necessity, current intent and correctly-open statusRoot Understand: deep pr-necessity / standard pr_common β†’ pr_context_auditsNecessity controller ladder; spending on superseded work without an explicit reason
Live UI applicability, branch/config/data truth and screenshotsFinal live-ui-validation β†’ live-ui-review β†’ live-ui-runtime; verified overlays own target/setup policy; root views used images onceSource fixes inside UI verification and automatic post-judgment fix tasks
Snapshot, discussion, pending-review deduplication and publicationRoot pr_snapshot / pr_common / review_delivery and publication skillWorker PR refetches; review authorship treated as edit or publish authority
Diagnosis, cleanup and PR-fix batchesRoot authorizes known production through k-diagnosing-bugs, k-knip, k-pr-fix-loop / pr_fix; SOP owns integrated final checks and scoped recoveryPrivate check loops, unrequested class-wide cleanup and automatic new-comment drains
Recall admission, corrections and durable learningRoot k-ai-kb β†’ smol-operator / cli; staged recall and one final learning batch; identical inline fallbackPer-turn scribes, leaf memory orchestration and dropping learning when delegation is forbidden
Convergence and prose alternativesExplicit k-converge invocation (declared exit + correctness filter) or requested k-text-tournament; root owns the enclosing lifecycleAutomatic convergence and repeated evaluator tournaments

Existing named profiles in Claude, Codex, Cursor, Copilot, Pi and OMP point to the same leaf contracts. Dynamic or unsupported adapters remain subject to the capability boundaries below; a contract reference does not prove profile discovery or runtime enforcement. Contract tests check required load edges, profile bindings and retained obligations. Inline source/evidence judgment remains necessary: text tests do not prove model compliance, lower token use or large-project quality parity.

Runtime controls and limits​

Coverage inventory​

Coverage is the join of the launcher, managed configs, provider wrappers and installed application surfaces. ,ai is not the harness inventory: it routes seven native CLIs and omits OMP and Crush. A native CLI check does not certify a subscription wrapper, local provider, hosted frontend or desktop application.

SurfaceConfigured scope and confidence boundary
Claude Code, Codex, Cursor CLI, Copilot CLINative profiles/hooks plus the separately listed provider routes. Native child controls and model selection are distinct claims.
Pi, OMPNative packages/extensions and role profiles; see controls below.
OpenCodeExisting work implementation workers; personal single-context. Strong category lanes are not configured.
Antigravity CLI (agy, registry key gemini)Native dynamic tiers and shared hooks. This is not evidence for a separate Gemini CLI or the desktop application.
CrushManaged crushrc loads the SOP. Installed v0.92.0 task defaults are read-only and share the large model; no managed category lanes or automatic memory hooks. Root-only native hooks do not certify child lifecycle enforcement.
OpenRouter wrappersClaude, Codex, Copilot and Cursor wrappers select the Pi backend matrix. Claude and Codex project managed native profiles with exact selectors, including distinct refute effort; the gate denies missing or stale role maps.
Codex subscription wrappers,claude-codex, ,copilot-codex, ,cursor-codex; use the Codex backend matrix, not the frontend's model catalog.
Copilot subscription wrappers,claude-copilot, ,codex-copilot, ,cursor-copilot; use the Copilot backend matrix and entitlement-filtered catalog.
llama.cpp wrappersClaude, Codex, Cursor and OpenCode launchers. A single local model is not a strong/cheap category ladder; native lifecycle controls still belong to the frontend.
,qDeliberately stripped Pi one-shot: no discovered extensions, skills, context or session persistence. It is not the staged coding workflow and does not inherit the Pi extension guards.
Cursor / Antigravity desktopSeparately installed application surfaces; CLI evidence does not establish desktop behavior.
tuicr / lgtmReview interfaces, not additional model orchestration harnesses.

Subscription adapters select child controls only through explicit @lane-<effort> selectors; a raw model ID never implies a worker role, even when it matches another registered lane. Raw root requests retain launch controls. Claude projects managed --agents definitions with exact registered pairs and preserves the profile body and skills. Selectors are stripped before the upstream request. The band hook removes call-level model overrides and denies missing/unavailable profiles, stale role maps, resume/fork requests, or conflicting controls. Projected tool lists exclude orchestration tools; root terminal-wakeup prevention remains a separate criterion.

,codex-copilot projects entitled lane selectors into a session-local Codex catalog and copies managed leaf profiles without native model/effort overrides. Codex applies profiles after spawn arguments, so both projections are required. The gate admits only projected roles and exact lanes; it rejects full-history forks, missing session projections, conflicting provider controls, and child-originated delegation. The profile copies retain their instructions and disable multi_agent. The session-local files live until the wrapper exits; native profiles remain unchanged. Relaunch sessions started before this projection existed.

The review runtime caveats in ~/.agents/skills/k-review/references/runtime-harnesses.md use this conditional support, not a blanket ,codex-copilot ban. Native child-tag transport and live provider completion are separate acceptance conditions: proving the former does not prove that a live refuter will complete.

Native child-tag transport is verified absent on ,copilot-codex: Copilot CLI 1.0.83 takes the child model from the per-subagent ~/.copilot/settings.json entry after the band hook rewrites the call's model, so an injected @lane-<effort> selector never reaches the loopback request (probed 2026-09-11). Delegation itself now runs on that route, unpinned, because the adapter restores the turn's output items on the streamed terminal event. ,cursor-codex and ,cursor-copilot remain unverified. Root sessions and native harness routes remain available. Fake-provider translation checks alone do not establish frontend selector acceptance.

Native Claude alias overrides are clamped in both directions: a cheaper alias must not weaken a strong role. Generic Codex/Copilot workers preserve a registered model/effort pair, not model membership alone. A model shared by multiple lane efforts requires an explicit matching effort. Cursor uses a model selector instead: its registered auto mechanical/memory rows intentionally omit effort. Backend-schema routes validate the backend contract, not the frontend's exception. Missing or malformed required lane data denies delegation on the configured mutating adapters. Copilot propagates those denials and rejects delegation when its band helper is missing, fails, or emits malformed output; optional memory errors remain independent.

Native controls​

  • All profile templates include the shared leaf contract. Former controller profiles are final judgment leaves, not nested controllers. Pi/OMP reviewer and blind-clarity profiles name the active root as their dispatcher; reviewers must not reload an ambient catalog or treat a controller-named sibling as an orchestrator.
  • Pi profiles set defaultContext: fresh, inheritProjectContext: false, inheritGlobalContext: false, inheritSkills: false, and child depth zero. This removes ambient instructions/catalog, not explicitly named role skills: pi-subagents 0.66.0 loads those separately in runs/foreground/execution.ts. The root supplies applicable project and safety constraints in the packet. Dispatch requires agentScope:"user" so project profiles cannot override managed models, tools, or inheritance. The runtime guard requires acceptance:false and rejects context/model/per-call skill overrides, composite workflows, host gates, scheduling and resume/steer recovery. Dispatch one ready packet per call; independent root packets may explicitly set async:true. Defaults are fresh and foreground. Native child identity suppresses root memory/reinforcement injection and blocks the subagent tool.
  • Claude generic implementation profiles omit agent tools. Unrestricted shell remains a limitation, not a sandbox guarantee.
  • OMP retains native depth pruning and disables background advisors. Explicit root async and effort controls remain available; implicit Bash/Eval backgrounding is disabled because it also affects blocking workers. Managed profiles (including native-name task, sonic, and scout shims) use blocking: true, no task tool, and no spawns allowlist; independent task batches still run concurrently before returning. Native reviewer shortcuts are disabled in favor of managed review profiles.
  • OMP's runtime extension latches the native yield capability as the leaf boundary and rejects leaf async, task, advisor and hub calls. This prevents those tool paths from creating work that outlives the return. Explicit-yield root sessions receive the same conservative restrictions. Ordinary roots retain named-process hub operations; peer sends remain blocked with native whitespace normalization. The native driver is not patched: an externally created pending job can still invalidate a yield. Do not confuse blocking:true with disabling background jobs or claim universal terminal enforcement.
  • OMP 18.1.14 passes context files, skills, and native child/Coop instructions through src/task/executor.ts; its profile parser does not expose Pi's inheritance flags. The managed before_agent_start hook replaces the leaf model prompt with its role, explicit packet/context/plan, native worktree restriction, and yield protocol/schema. It removes the inherited SOP/catalog and native private-QA/peer-wakeup instructions; named skill autoload messages remain. Unknown, ambiguous, or unmanaged prompt frames abort before provider dispatch instead of falling back to a controller prompt. This controls the model-facing prompt, not physical SDK discovery, native revival, or later third-party prompt rewrites. Do not use an adapter unattended when it cannot enforce the required no-orchestration/terminal boundary.
  • Former controller leaves omit declared edit/write/agent/task tools where supported. Cursor retains readonly: false for its existing shell/MCP access caveat; prompt-level read-only instructions are not runtime write isolation.
  • Shared startup/per-turn hooks skip Codex/Claude child payloads carrying agent_id before topic lookup or root-context injection. Root hooks retain filtered KB retrieval and staging; only the root owns admission and final learning. No per-turn scribe or automatic convergence. Topic binding, worklogs, context-disable sentinels, and reinforcement remain.
  • Codex disables features.multi_agent in every managed child role, including native default, worker, and explorer overrides. Generic roles leave model selection to the existing band hook, preserving explicit registry-backed lane picks; dedicated roles retain their declared bands. Root collaboration remains enabled. This is CLI role configuration, not proof that a hosted frontend honors local role files.
  • OpenCode work retains its existing implementation workers/models, now with the shared leaf prompt and native permission.task: deny. Only the work root may dispatch worker_*; personal remains single-context. The newly available built-in scout is disabled alongside existing built-ins. The memory plugin resolves native parentID: root recall/learning plumbing remains, children keep worklogs without root recall, and child/unknown-identity task calls fail closed. Existing worker-session resume requests are rejected. OpenCode still has no category-backed strong research/review/refutation adapter; do not substitute its implementation workers for those stages.
  • Cursor/Copilot retain their model-band adapters. Claude/Copilot managed profile tool allowlists omit delegation. Cursor CLI does not discover user custom profiles, and Antigravity's native dynamic tiers are not leaf lifecycle enforcement. These surfaces still require the root's packet/no-resume discipline; do not run unattended isolated work where the runtime boundary is unavailable. Unrestricted shell/MCP is not an agent sandbox in any harness.

OMP 18.1.14 dispatches blocking profiles through its synchronous fan-out path (src/task/index.ts); src/discovery/helpers.ts parses the flags. The root guard uses native pi.pi.discoverAgents(cwd) before dispatch. Task packets must name their profiles explicitly, including every batch item; implicit native defaults are rejected. Each selected definition must resolve to its enabled user-profile file under the canonical agent directory and carry blocking, explicit model/tool and leaf-boundary controls. Project, plugin-path and bundled replacements are rejected. Settings-level model overrides, worker advisors and prewalk handoffs are rejected too; explicit off values and final refute profiles using @advisor remain available. Eval can select profiles dynamically, so its admission checks every enabled definition. This can block non-agent Eval code when an enabled definition is unsafe; use ordinary root tools or correct the profile, not an unguarded dispatcher.

OMP CLI task/Eval dispatch also checks effective Bash/Eval foreground settings immediately before dispatch. Project/CLI overlays can override the managed file; unsafe or unavailable values block dispatch with the offending keys. The guard reads the CLI's bundled pi.pi.Settings.instance and reloads disk layers through the same native method as subagent preflight; it sets no values or overrides. A concurrent terminal-marker failure also blocks an admission awaiting discovery. Source/legacy-module imports can resolve a different singleton and silently return defaults in the shipped npm CLI. This is a CLI-root control, not certification of secondary roots created with isolated SDK settings. Those roots remain unsupported for unattended dispatch.

Eval has a separate native bypass. In OMP 18.1.14, eval/js/tool-bridge.ts routes __agent__, __workpool__ and __completion__ before ordinary tool lookup and wrapping; both Eval runtimes use that bridge. agent() registers background work with keep-alive execution in eval/agent-bridge.ts, irrespective of a profile's blocking flag. workpool() creates persistent pools; completion() resolves its own tier/effort and starts a model request without a managed profile. Passing the outer Eval admission guard does not enforce the inner helper's lifecycle or model lane. These helpers and their synthetic bridges are prohibited as managed dispatch routes; use explicit native task packets instead. Ordinary Eval is not disabled, and no runtime interception of these helper calls is claimed.

The k-omp adapter now puts task dispatch and named-process hub operations under its root-only section. A leaf skips that section rather than receiving a generic instruction to perform agent handoffs. The root-moves instruction scanner recognizes the quoted native task command as well as β€œTask tool”; its counterexamples retain prohibitions, ordinary task-state prose and root-gated dispatch.

Native OMP plan-mode dispatch is a confirmed unsupported path: structured-subagent.ts sets restrictToolNames, and sdk.ts loads an empty extension list for those children. That skips the managed leaf prompt/tool/terminal hooks. A per-call Eval-defined tools list alone does not set this restriction. The public extension context has no plan-mode getter, while ACP and plan-yolo can change mode without the TUI's transcript entries; transcript inspection would not close the gap. Do not dispatch unattended workers in native plan mode or restricted SDK sessions. This restriction does not change the strong @plan model role or prohibit root planning. Profile admission is not a fix for extension omission, concurrent external file edits, or altered files already inside the user profile directory.

A native yield while owned asynchronous jobs remain is provisional: the executor may collect results and request another yield before returning to the parent. A test observing two yields before the parent receives a result proves additional internal turns, not post-completion resurrection. Terminal tests must observe the parent-visible result and subsequent delivery separately; neither a one-yield fixture nor the foreground-settings guard proves every native revival path is sealed.

The OMP root also consumes the native task:subagent:lifecycle finalized event. Only completed, failed, or aborted creates an exclusive <sessionFile>.k-leaf-terminal marker beside that worker's native transcript. The native finalizer emits this after saving and classifying the result, not at a provisional yield. The marker leaves the transcript and result untouched. Managed leaves check it before prompts, provider requests and tool execution, including a cold reopen of the same session. This prevents guarded execution after the observed final event; it does not prevent native registry retention or session construction. Leaves without persisted session identity are blocked. A failed marker write makes the root task/Eval receipt an error and blocks further dispatch; other completed workers are still sealed.

Do not treat pre-existing workers without markers, removed/relocated markers, disabled/replaced extensions, or arbitrary shell/MCP as covered by that boundary. Do not resume a completed workpool batch as a new packet. The native-pipeline fixture uses a synthetic leaf identity and finalized event, so full task/revival confidence still requires a separately authorized native worker pilot.

OpenCode's task/no-resume guards register even when optional memory helper files are absent. Each available memory callback remains active independently; missing recall or worklog plumbing does not turn off lifecycle controls.

Prompt contracts do not prove native enforcement. Unsupported autonomous lifecycle controls require a visible capability limitation rather than a claimed guarantee. No universal spend cap is asserted; usage accounting must include children/advisors when the harness exposes it. The policy evaluation matrix includes every capability-snapshot harness, including OMP and Crush; runtime-only child bindings are reported unsupported for static-pinning scenarios, not silently omitted.

These controls prevent specific sources of extra work; they do not classify arbitrary shell commands as research, production or verification. The root still owns the single final check plan. A command allowlist broad enough for implementation is not proof that a worker cannot run private QA. Runtime boundary tests and a small worker pilot also do not establish better quality or speed on a matched Kibana-scale workload.

Sources​

Core: home/readonly_AGENTS.md Β§Β§3.5–3.7. Leaf: home/dot_config/exact_tmux/agent_prompts/leaf-boundary.txt. Recipes: home/exact_dot_agents/exact_skills/. Runtime: shared hooks plus Pi/OMP extensions and profile templates. See model tiering, spec/build, and cross-agent memory.

Native configuration references: Codex custom agents β€œloads these files as configuration layers for spawned sessions”; OpenCode task permissions β€œControl which subagents an agent can invoke via the Task tool with permission.task.” Installed source remains authoritative for the tested runtime version.

Crush source: v0.92.0 hook scope: β€œSub-agents (the agent task tool, agentic_fetch, etc.) run without hook wrapping on their internal tools.”