Skip to main content

Creation workflow

The creation-side counterpart to the review workflow applies the same rigor primitives to building things instead of judging them: evidence-gated phases, fixed return shapes, adversarial verification, and a completion gate.

The steering model is two human gates: approve the contract before execution, then read the report after it. Everything between runs hands-free.

Ordinary freeform implementation does not have to enter this formal flow. Verification stays inline by default; k-proof and ,proof add a smaller repo-external receipt only for an explicit receipt request, an auditable risky effect, or a named handoff/resume consumer.

The two artifacts (memory vs contract)โ€‹

ArtifactRoleMutation rule
/tmp/specs/<pwd>/<topic>.txtConversation memory: what we currently believe the user wants. Hook-injected at session start.Rewritten freely as the intent loop converges. Allowed to be wrong.

The intent spec remembers the discussion; the packet is a signed order.

The packet is never a mechanical transform of the intent spec. Nothing enters it on the .txt's word alone: every criterion is re-derived from evidence and its check is run once, observed red, before it may appear.

The k-spec skill writes a packet: pointer line into <topic>.txt so session-start injection tells a fresh session the contract exists.

On default branches, k-spec requires a named topic first because the session-<id> fallback would strand the packet.

Two pivots keep the contract honest in both directions: empirical forks route out to k-prototype, where the verdict returns to the packet; mid-build premise contradictions route back to gate 1. Both โ€” and all flow-to-flow movement โ€” are mapped in Choose your flow.

Using itโ€‹

Lifecycle:

idea/issue
โ””โ”€ spec skill: necessity check โ†’ fork-closing interview (interview-me discipline,
empirical forks โ†’ prototype) โ†’ acceptance criteria with run-once RED checks
โ†’ packet written + shown
โ””โ”€ [HUMAN GATE 1: approve packet]
โ”œโ”€ /k-build ............ in-session hands-free implementation
โ”œโ”€ k-compose-issue ... publishable issue text + publication packet
โ””โ”€ review (plan mode) adversarial review of the packet itself
โ””โ”€ [HUMAN GATE 2: read the report]

/k-build phase topologyโ€‹

PhaseOwnerGate
1. Spec gatecontroller + humanpacket exists, checks red-proven; explicit approval
2. Plancontrollerper-step verification defined; Ownership Gate over touched paths
3. Executecontrollercriteria ledger updated per step; never past a red step; ยง3.3 reset on 2ร—
4. Mechanical gatescontrollerrepo lint/type/tests discovered, run, looped to green
5. Live-UI proofk-ui-proof (inline)visual criteria verified head-only against the built runtime; each proof set captured to its own distinct /tmp/<folder-name>/ and opened/provided
6. Adversarial verifycriteria-verifier lanechecks re-run from clean tree; refutation verdicts + scope audit
7. Post-review stagecontrollerfour dimensions over the implementation diff
8. Reportcontroller + humanmandated output block; completion gate

The criteria ledger is the run's spine. It has one row per acceptance criterion (red / green / judgment-met / judgment-unmet / blocked), each with command-level evidence, plus a verification verdict (confirmed / refuted / undecidable).

Verdicts are evidence, not decisions. The controller flips a row only after checking the refutation addresses the row's actual claim.

The criteria-verifier laneโ€‹

Worker contract: k-build/references/criteria-verifier.md.

The lane owns refutation order (claim truth โ†’ criterion truth โ†’ reachability โ†’ durability), a scope audit against the packet's binding out-of-scope list, and missing-criteria candidates.

Per-harness profiles are rendered from the same agent_review_models registry the review verifier uses. Cursor, Copilot, Gemini, Codex, and Pi ship a criteria-verifier profile.

Claude runs the lane degraded on the session model with refutation framing, reported as families=same (degraded). This mirrors the adversarial-verifier convention in Cross-harness subagents.

Live-UI proof (phase 5)โ€‹

When any acceptance criterion's evidence is visual โ€” a judgment: criterion naming a screenshot/visual comparison, or an in-scope UI-facing change with a stated visual goal โ€” /k-build runs the k-ui-proof skill.

It is the creation-side sibling of the review flow's live-ui-review: same runtime machinery, opposite direction. live-ui-review compares PR/head against base to find regressions; k-ui-proof verifies the built runtime head-only against its intended visual and captures the screenshot set that proves it.

Both share one mode-neutral contract โ€” k-agent-review/references/live-ui-runtime.md โ€” for target-packet resolution, Playwriter preflight, readiness, runtime start, the data/setup ladder, screenshot artifacts, and the runtime safety boundary.

Each mode file adds only its oracle, comparison model, and return shape.

k-ui-proof runs inline in /k-build, which already holds Playwriter and local/dev mutation permissions, so it needs no isolated subagent profile.

It returns a per-criterion met / unmet / blocked verdict. The controller sets the ledger's judgment-met/judgment-unmet row from it; an unmet returns to phase 3 like a red step.

The controller reports the screenshot manifest. Each screenshot/pair/set lives in its own distinct /tmp/<folder-name>/ folder, so k-compose-pr can upload and embed the shots.

Windows/VirtualBox coverage is a separate manual skill, k-live-ui-windows, connecting Playwriter to a guest browser over CDP through a host NAT port-forward. It is never auto-triggered by either mode; load it by hand only when the user explicitly asks for Windows/VirtualBox verification this turn.

A verify failure sends bounded evidence back to implementation, retries up to the configured budget, and then reports a blocker for human decision.