Orchestrate plans from context to merged PRs with Herdr #30

Closed
opened 2026-10-06 12:36:53 -04:00 by steve · 2 comments
Owner

Problem Statement

Turning a plan into completed, reviewed, merged work currently requires an operator or agent to coordinate several skills, issue-tracker operations, code workers, reviewers, Git worktrees, and Herdr sessions by hand. The workflow is easy to lose across sessions and machines: agents may be duplicated after an uncertain outcome, tasks may be started before prerequisites are ready, and recovery depends on reconstructing state from transient conversations.

The user needs one explicit orchestration entry point that accepts a plan from the current conversation or an @file, then carries it through specification, tracker tickets, implementation, review, and merge. It must work whether the user invokes it from the same machine that runs the agents or from a separate orchestrator system connected to a selected worker node.

Solution

Provide an explicitly invoked Pi skill that uses Herdr and Pi agents—without a Bun runtime or separate orchestration CLI—to turn a plan into a published spec and child tickets, then drive one task ticket at a time through implementation, four-axis review, and merge. The user selects the worker node at invocation; it may be local or remote. The worker node is authoritative for run state, allowing a later orchestrator system to reconnect and sync state before resuming.

The workflow composes the existing specification, ticketing, worktree, Herdr, review, and PR skills. It persists enough state and handoffs on the worker node to reconcile side effects and recover safely. It stops for human judgment instead of guessing, and only reports completion when the agreed acceptance criteria are verified and all in-scope tasks are complete.

User Stories

  1. As a developer, I want to explicitly invoke one orchestration skill with a goal, so that a multi-session plan can progress without manually coordinating every agent.
  2. As a developer, I want to provide a plan in the current conversation, so that I do not need to first save it to a file.
  3. As a developer, I want to reference a plan with @file, so that an existing document can be the orchestration input.
  4. As a developer, I want the orchestrator to create and publish a spec from the supplied plan, so that the goal and acceptance criteria have a durable tracker record.
  5. As a developer, I want the spec broken into linked, dependency-aware task tickets, so that work is discoverable and can be resumed from the tracker.
  6. As a developer, I want task tickets sized as complete, verifiable slices, so that workers can deliver useful outcomes without relying on hidden context.
  7. As a developer, I want to select a worker node at invocation, so that execution happens on the machine with the required repository, tools, and credentials.
  8. As a developer, I want the worker node to be allowed to be either the orchestrator system or a remote machine, so that local and remote execution use the same workflow.
  9. As an operator, I want clear role terminology—orchestrator system for the invoking machine, worker node for the selected execution machine, and worker agent for a subagent—so that instructions and recovery records are unambiguous.
  10. As an operator, I want Herdr to discover sessions, workspaces, and panes rather than guessing identifiers, so that the orchestrator does not target the wrong agent or repository.
  11. As an operator, I want the selected pane and repository/ref checked before dispatch, so that work starts in the intended environment.
  12. As a developer, I want the orchestrator to execute one tracker task ticket at a time, so that scope and review remain attributable and recoverable.
  13. As a developer, I want each implementation task to run in an isolated Git worktree and branch, so that worker changes do not collide with the base checkout or sibling work.
  14. As a developer, I want workers to receive the full ticket, acceptance criteria, scope boundaries, and required upstream handoffs, so that they can work without access to another agent's conversation.
  15. As a developer, I want independent verification for meaningful behavior changes, so that a worker's self-report is not the only evidence of acceptance.
  16. As a developer, I want reviews to cover correctness/spec and standards through both the dual-review and code-review workflows, so that all four required review axes pass before merge.
  17. As a developer, I want actionable review findings sent to a scoped fix worker and rechecked, so that review feedback can be resolved without restarting the entire plan.
  18. As a developer, I want review/fix rounds capped at three, so that unresolved findings eventually stop for human judgment.
  19. As an operator, I want transient infrastructure failures retried at most twice after the first attempt, so that temporary failures can recover without unbounded agent churn.
  20. As an operator, I want semantic failures, failed checks, and rejected reviews not blindly retried, so that the orchestrator addresses causes instead of repeating the same failed action.
  21. As an operator, I want uncertain side effects reconciled before retrying, so that recovery does not create duplicate agents, worktrees, comments, PRs, merges, or state transitions.
  22. As an operator, I want authentication and permission failures, exhausted retry limits, unresolved state mismatches, and decisions requiring judgment escalated to a human, so that the orchestrator does not guess past a blocker.
  23. As a developer, I want task PRs created and linked to their tracker tickets, so that implementation and review evidence stay attached to the work.
  24. As a developer, I want merges blocked until all four review axes pass, required local checks pass, available CI is green, the target branch is still correct, and no human decision remains, so that incomplete work is not integrated.
  25. As a developer, I want merged task tickets closed with a recorded outcome, so that tracker status reflects completed work.
  26. As a developer, I want the parent spec ticket closed only after every in-scope task is complete, merged, verified, and run state is safely persisted, so that a closed spec means the workflow actually finished.
  27. As an operator, I want the worker node to hold the authoritative run state, so that a different orchestrator system can reconnect to the same node and resume.
  28. As an operator, I want the orchestrator system to sync state from the worker node before resuming, so that it does not act on a stale local copy.
  29. As an operator, I want run state to include task progress, attempts, branches/worktrees, Herdr identifiers, handoffs, and operator decisions, so that a restarted planner can reconcile rather than duplicate work.
  30. As an operator, I want worker handoffs retained verbatim with task and execution metadata, so that downstream agents and humans can rely on the original evidence.
  31. As an operator, I want clean task worktrees removed and only run-owned Herdr panes closed after completion, so that completed runs do not leave unnecessary resources behind.
  32. As an operator, I want dirty or unresolved worktrees preserved, so that cleanup cannot destroy unfinished work or recovery evidence.
  33. As an operator, I want the worker node's canonical checkout, orchestration state, handoffs, and logs retained, so that later inspection or recovery remains possible.
  34. As an operator, I want retained state deleted only by explicit user action, so that cleanup never silently removes the recovery record.
  35. As a developer, I want human-interaction PRs marked for human review and resumable within the linked ticket/PR scope, so that automation respects the established review handoff.
  36. As a developer, I want the orchestrator to report merged work, PR links, verification evidence, human handoffs, blockers, and remaining work, so that completion status is actionable and truthful.

Implementation Decisions

  • Implement the workflow as a Pi skill and agent instructions that invoke Herdr directly. Do not add a Bun runtime, standalone orchestration CLI, or parallel process-control framework.
  • Use the existing specification, ticketing, Herdr, worktree, dual-review, code-review, PR, and forge skills for their specialized work. Make to-spec, to-tickets, and dual-review invocable by agents as required by the orchestration workflow.
  • Keep roles distinct: the orchestrator system invokes and observes the workflow; the user-selected worker node runs the planner and worker agents and owns repository work; worker agents implement or verify scoped tasks. The worker node may be local or remote and is selected at skill invocation.
  • Use Herdr to discover and control the selected worker node's sessions, workspaces, panes, and worktrees. Inspect the exact target and repository/ref before dispatch. Never infer pane IDs, target a pane by display label alone, or modify unrelated panes.
  • Accept either the current conversation context or an @file reference as the plan input. Publish the generated spec to the current repository's issue tracker, then create typed/linked task tickets with blocker relationships and acceptance criteria.
  • Follow the repo's current Gitea workflow. Apply ready-for-agent to the published spec and agent-grabbable task tickets. Use needs-review for human-interaction PRs and ready-for-agent to resume only within the spec's linked ticket/PR scope.
  • Process one tracker task ticket at a time. Select only in-scope, open, unblocked work; claim it before implementation and re-query tracker state as work advances.
  • Give each implementation task an isolated worktree and branch. Use the current repository checkout when the selected worker node is local; when remote, use the canonical .bare worktree layout and the agreed owner/repository worktree location.
  • Planners coordinate and verify state; workers implement; independent verifiers report evidence. Planners do not edit product code or merge. Workers do not merge unless assigned an explicit merge task.
  • Persist the run's plan, task state, handoffs, and operator/recovery log under the run-scoped .orchestrate/<slug> directory in the worker node's repository checkout. The worker node is the source of truth. A new orchestrator system reconnects to that node and syncs from it before resuming; a local copy is not authoritative. Do not create a separate Git branch solely to transfer this state.
  • Persist state immediately after side effects and record enough Herdr, branch, worktree, and task identity to reconcile execution after interruption. On restart, inspect persisted state and native Herdr status before creating or restarting agents. Never duplicate work solely because the planner session restarted.
  • Retry transient infrastructure failures only, with at most two retries after the first attempt (three attempts total). Keep this retry budget separate from the three review/fix rounds. Do not automatically retry semantic failures, failed checks, or rejected reviews without correcting the cause.
  • For uncertain side effects, reconcile the tracker, Git, worker node, or Herdr state before repeating an operation. If the outcome remains uncertain, stop and escalate; do not blindly replay non-idempotent operations.
  • Escalate immediately for authentication/permission failures, unresolved state mismatches, exhausted retry/review limits, unsafe ambiguity, or decisions requiring human judgment. Preserve state and evidence at every escalation.
  • Merge only when required reviews and local checks pass, CI is green when available, the target branch remains the expected branch, and no human decision or blocker remains. Do not build CI as part of this feature.
  • Close task tickets after their PRs are merged and completion is verified. Close the parent spec ticket only when every in-scope task is complete, merged, verified, and run state is safely persisted. Leave it open when any work is blocked, failed, or unresolved.
  • At completion, remove clean task worktrees and close only Herdr panes created by this run. Preserve dirty or unresolved worktrees. Retain the worker node's canonical checkout, run state, handoffs, and logs; never delete retained state automatically. Cleanup of retained data requires an explicit user action.
  • Keep the existing Cursor /orchestrate plugin/backend unchanged.

Testing Decisions

  • A good test exercises the user-visible orchestration outcome, not prompt wording, internal role decomposition, or implementation-specific helper structure. It provides evidence that tracker state, Git worktrees/branches, Herdr agents, recovery, review gates, and cleanup match the accepted contract.
  • Use one manual end-to-end acceptance seam at the explicit Pi skill invocation in a disposable repository and named Herdr session. Run actual Herdr and Pi agents; do not introduce a separate runtime test harness.
  • The end-to-end check covers plan input from both conversation and @file, spec and ticket publication, task claim/implementation, isolated worktree use, handoff and verification, four-axis review, PR/merge gates, persisted worker-node state, retry/escalation behavior, and terminal cleanup.
  • Include recovery by starting from a different orchestrator system, reconnecting to the same worker node, syncing state, reconciling Herdr/tracker state, and resuming without duplicate agents or side effects.
  • Exercise both local and remote worker-node selection, including a clear stop when the selected node/session/repository is missing, ambiguous, or not ready.
  • Exercise transient failure and retry exhaustion, semantic failure with no retry, uncertain side-effect reconciliation, blocked/failed task preservation, a review requiring fixes, and a human-interaction handoff.
  • Verify only run-owned Herdr panes and clean worktrees are cleaned up; unrelated panes and dirty/unresolved worktrees remain untouched; retained state is not automatically deleted.
  • This repository is currently a Markdown skills repository with no runtime package or established automated test harness. Existing prior art is the operational guidance in the Herdr orchestration and implementation-orchestrator skills. The accepted validation seam is a real, isolated manual workflow rather than a new Bun CLI or test framework.

Out of Scope

  • A Bun runtime, standalone orchestration CLI, daemon, or new process-control framework.
  • Changing or replacing the existing Cursor /orchestrate implementation.
  • Building the forge's CI system; require CI when available, but the user handles introducing it separately.
  • Managing tracker, Git, Herdr, or provider credentials, authentication, or login sessions.
  • Running multiple tracker task tickets concurrently.
  • Automatically deleting retained orchestration state, logs, or the canonical worker checkout.
  • Operating on unrelated tracker tickets, repositories, panes, branches, worktrees, or PRs.
  • Merging when review, verification, CI, scope, or human-approval gates are unresolved.

Further Notes

## Problem Statement Turning a plan into completed, reviewed, merged work currently requires an operator or agent to coordinate several skills, issue-tracker operations, code workers, reviewers, Git worktrees, and Herdr sessions by hand. The workflow is easy to lose across sessions and machines: agents may be duplicated after an uncertain outcome, tasks may be started before prerequisites are ready, and recovery depends on reconstructing state from transient conversations. The user needs one explicit orchestration entry point that accepts a plan from the current conversation or an `@file`, then carries it through specification, tracker tickets, implementation, review, and merge. It must work whether the user invokes it from the same machine that runs the agents or from a separate orchestrator system connected to a selected worker node. ## Solution Provide an explicitly invoked Pi skill that uses Herdr and Pi agents—without a Bun runtime or separate orchestration CLI—to turn a plan into a published spec and child tickets, then drive one task ticket at a time through implementation, four-axis review, and merge. The user selects the worker node at invocation; it may be local or remote. The worker node is authoritative for run state, allowing a later orchestrator system to reconnect and sync state before resuming. The workflow composes the existing specification, ticketing, worktree, Herdr, review, and PR skills. It persists enough state and handoffs on the worker node to reconcile side effects and recover safely. It stops for human judgment instead of guessing, and only reports completion when the agreed acceptance criteria are verified and all in-scope tasks are complete. ## User Stories 1. As a developer, I want to explicitly invoke one orchestration skill with a goal, so that a multi-session plan can progress without manually coordinating every agent. 2. As a developer, I want to provide a plan in the current conversation, so that I do not need to first save it to a file. 3. As a developer, I want to reference a plan with `@file`, so that an existing document can be the orchestration input. 4. As a developer, I want the orchestrator to create and publish a spec from the supplied plan, so that the goal and acceptance criteria have a durable tracker record. 5. As a developer, I want the spec broken into linked, dependency-aware task tickets, so that work is discoverable and can be resumed from the tracker. 6. As a developer, I want task tickets sized as complete, verifiable slices, so that workers can deliver useful outcomes without relying on hidden context. 7. As a developer, I want to select a worker node at invocation, so that execution happens on the machine with the required repository, tools, and credentials. 8. As a developer, I want the worker node to be allowed to be either the orchestrator system or a remote machine, so that local and remote execution use the same workflow. 9. As an operator, I want clear role terminology—**orchestrator system** for the invoking machine, **worker node** for the selected execution machine, and **worker agent** for a subagent—so that instructions and recovery records are unambiguous. 10. As an operator, I want Herdr to discover sessions, workspaces, and panes rather than guessing identifiers, so that the orchestrator does not target the wrong agent or repository. 11. As an operator, I want the selected pane and repository/ref checked before dispatch, so that work starts in the intended environment. 12. As a developer, I want the orchestrator to execute one tracker task ticket at a time, so that scope and review remain attributable and recoverable. 13. As a developer, I want each implementation task to run in an isolated Git worktree and branch, so that worker changes do not collide with the base checkout or sibling work. 14. As a developer, I want workers to receive the full ticket, acceptance criteria, scope boundaries, and required upstream handoffs, so that they can work without access to another agent's conversation. 15. As a developer, I want independent verification for meaningful behavior changes, so that a worker's self-report is not the only evidence of acceptance. 16. As a developer, I want reviews to cover correctness/spec and standards through both the dual-review and code-review workflows, so that all four required review axes pass before merge. 17. As a developer, I want actionable review findings sent to a scoped fix worker and rechecked, so that review feedback can be resolved without restarting the entire plan. 18. As a developer, I want review/fix rounds capped at three, so that unresolved findings eventually stop for human judgment. 19. As an operator, I want transient infrastructure failures retried at most twice after the first attempt, so that temporary failures can recover without unbounded agent churn. 20. As an operator, I want semantic failures, failed checks, and rejected reviews not blindly retried, so that the orchestrator addresses causes instead of repeating the same failed action. 21. As an operator, I want uncertain side effects reconciled before retrying, so that recovery does not create duplicate agents, worktrees, comments, PRs, merges, or state transitions. 22. As an operator, I want authentication and permission failures, exhausted retry limits, unresolved state mismatches, and decisions requiring judgment escalated to a human, so that the orchestrator does not guess past a blocker. 23. As a developer, I want task PRs created and linked to their tracker tickets, so that implementation and review evidence stay attached to the work. 24. As a developer, I want merges blocked until all four review axes pass, required local checks pass, available CI is green, the target branch is still correct, and no human decision remains, so that incomplete work is not integrated. 25. As a developer, I want merged task tickets closed with a recorded outcome, so that tracker status reflects completed work. 26. As a developer, I want the parent spec ticket closed only after every in-scope task is complete, merged, verified, and run state is safely persisted, so that a closed spec means the workflow actually finished. 27. As an operator, I want the worker node to hold the authoritative run state, so that a different orchestrator system can reconnect to the same node and resume. 28. As an operator, I want the orchestrator system to sync state from the worker node before resuming, so that it does not act on a stale local copy. 29. As an operator, I want run state to include task progress, attempts, branches/worktrees, Herdr identifiers, handoffs, and operator decisions, so that a restarted planner can reconcile rather than duplicate work. 30. As an operator, I want worker handoffs retained verbatim with task and execution metadata, so that downstream agents and humans can rely on the original evidence. 31. As an operator, I want clean task worktrees removed and only run-owned Herdr panes closed after completion, so that completed runs do not leave unnecessary resources behind. 32. As an operator, I want dirty or unresolved worktrees preserved, so that cleanup cannot destroy unfinished work or recovery evidence. 33. As an operator, I want the worker node's canonical checkout, orchestration state, handoffs, and logs retained, so that later inspection or recovery remains possible. 34. As an operator, I want retained state deleted only by explicit user action, so that cleanup never silently removes the recovery record. 35. As a developer, I want human-interaction PRs marked for human review and resumable within the linked ticket/PR scope, so that automation respects the established review handoff. 36. As a developer, I want the orchestrator to report merged work, PR links, verification evidence, human handoffs, blockers, and remaining work, so that completion status is actionable and truthful. ## Implementation Decisions - Implement the workflow as a Pi skill and agent instructions that invoke Herdr directly. Do not add a Bun runtime, standalone orchestration CLI, or parallel process-control framework. - Use the existing specification, ticketing, Herdr, worktree, dual-review, code-review, PR, and forge skills for their specialized work. Make `to-spec`, `to-tickets`, and `dual-review` invocable by agents as required by the orchestration workflow. - Keep roles distinct: the orchestrator system invokes and observes the workflow; the user-selected worker node runs the planner and worker agents and owns repository work; worker agents implement or verify scoped tasks. The worker node may be local or remote and is selected at skill invocation. - Use Herdr to discover and control the selected worker node's sessions, workspaces, panes, and worktrees. Inspect the exact target and repository/ref before dispatch. Never infer pane IDs, target a pane by display label alone, or modify unrelated panes. - Accept either the current conversation context or an `@file` reference as the plan input. Publish the generated spec to the current repository's issue tracker, then create typed/linked task tickets with blocker relationships and acceptance criteria. - Follow the repo's current Gitea workflow. Apply `ready-for-agent` to the published spec and agent-grabbable task tickets. Use `needs-review` for human-interaction PRs and `ready-for-agent` to resume only within the spec's linked ticket/PR scope. - Process one tracker task ticket at a time. Select only in-scope, open, unblocked work; claim it before implementation and re-query tracker state as work advances. - Give each implementation task an isolated worktree and branch. Use the current repository checkout when the selected worker node is local; when remote, use the canonical `.bare` worktree layout and the agreed owner/repository worktree location. - Planners coordinate and verify state; workers implement; independent verifiers report evidence. Planners do not edit product code or merge. Workers do not merge unless assigned an explicit merge task. - Persist the run's plan, task state, handoffs, and operator/recovery log under the run-scoped `.orchestrate/<slug>` directory in the worker node's repository checkout. The worker node is the source of truth. A new orchestrator system reconnects to that node and syncs from it before resuming; a local copy is not authoritative. Do not create a separate Git branch solely to transfer this state. - Persist state immediately after side effects and record enough Herdr, branch, worktree, and task identity to reconcile execution after interruption. On restart, inspect persisted state and native Herdr status before creating or restarting agents. Never duplicate work solely because the planner session restarted. - Retry transient infrastructure failures only, with at most two retries after the first attempt (three attempts total). Keep this retry budget separate from the three review/fix rounds. Do not automatically retry semantic failures, failed checks, or rejected reviews without correcting the cause. - For uncertain side effects, reconcile the tracker, Git, worker node, or Herdr state before repeating an operation. If the outcome remains uncertain, stop and escalate; do not blindly replay non-idempotent operations. - Escalate immediately for authentication/permission failures, unresolved state mismatches, exhausted retry/review limits, unsafe ambiguity, or decisions requiring human judgment. Preserve state and evidence at every escalation. - Merge only when required reviews and local checks pass, CI is green when available, the target branch remains the expected branch, and no human decision or blocker remains. Do not build CI as part of this feature. - Close task tickets after their PRs are merged and completion is verified. Close the parent spec ticket only when every in-scope task is complete, merged, verified, and run state is safely persisted. Leave it open when any work is blocked, failed, or unresolved. - At completion, remove clean task worktrees and close only Herdr panes created by this run. Preserve dirty or unresolved worktrees. Retain the worker node's canonical checkout, run state, handoffs, and logs; never delete retained state automatically. Cleanup of retained data requires an explicit user action. - Keep the existing Cursor `/orchestrate` plugin/backend unchanged. ## Testing Decisions - A good test exercises the user-visible orchestration outcome, not prompt wording, internal role decomposition, or implementation-specific helper structure. It provides evidence that tracker state, Git worktrees/branches, Herdr agents, recovery, review gates, and cleanup match the accepted contract. - Use one manual end-to-end acceptance seam at the explicit Pi skill invocation in a disposable repository and named Herdr session. Run actual Herdr and Pi agents; do not introduce a separate runtime test harness. - The end-to-end check covers plan input from both conversation and `@file`, spec and ticket publication, task claim/implementation, isolated worktree use, handoff and verification, four-axis review, PR/merge gates, persisted worker-node state, retry/escalation behavior, and terminal cleanup. - Include recovery by starting from a different orchestrator system, reconnecting to the same worker node, syncing state, reconciling Herdr/tracker state, and resuming without duplicate agents or side effects. - Exercise both local and remote worker-node selection, including a clear stop when the selected node/session/repository is missing, ambiguous, or not ready. - Exercise transient failure and retry exhaustion, semantic failure with no retry, uncertain side-effect reconciliation, blocked/failed task preservation, a review requiring fixes, and a human-interaction handoff. - Verify only run-owned Herdr panes and clean worktrees are cleaned up; unrelated panes and dirty/unresolved worktrees remain untouched; retained state is not automatically deleted. - This repository is currently a Markdown skills repository with no runtime package or established automated test harness. Existing prior art is the operational guidance in the Herdr orchestration and implementation-orchestrator skills. The accepted validation seam is a real, isolated manual workflow rather than a new Bun CLI or test framework. ## Out of Scope - A Bun runtime, standalone orchestration CLI, daemon, or new process-control framework. - Changing or replacing the existing Cursor `/orchestrate` implementation. - Building the forge's CI system; require CI when available, but the user handles introducing it separately. - Managing tracker, Git, Herdr, or provider credentials, authentication, or login sessions. - Running multiple tracker task tickets concurrently. - Automatically deleting retained orchestration state, logs, or the canonical worker checkout. - Operating on unrelated tracker tickets, repositories, panes, branches, worktrees, or PRs. - Merging when review, verification, CI, scope, or human-approval gates are unresolved. ## Further Notes - Domain terms: **orchestrator system** is the system where the skill is invoked; **worker node** is the user-selected machine that runs agents and repository work, local or remote; **worker agent** is a subagent running on that node. - The worker node is authoritative for run state. Recovery from another orchestrator system assumes access to the same worker node and its retained repository checkout. - The retry and lifecycle decisions were resolved in [Orchestrate plans through implementation and PR completion](http://gitea.sagacity.ca/steve/skills/issues/27): [Set bounded retries and human escalation policy](http://gitea.sagacity.ca/steve/skills/issues/28) and [Define terminal cleanup and parent-spec lifecycle](http://gitea.sagacity.ca/steve/skills/issues/29). - The current Gitea host has no CI; CI gates apply when CI becomes available. Do not expand this specification to build CI.
steve added the ready-for-agent label 2026-10-06 12:36:53 -04:00
steve added in-progress and removed ready-for-agent labels 2026-10-06 12:40:31 -04:00
Author
Owner

PR opened: #31

PR opened: http://gitea.sagacity.ca/steve/skills/pulls/31
steve added needs-review and removed in-progress labels 2026-10-06 12:52:54 -04:00
steve removed the needs-review label 2026-10-06 13:34:20 -04:00
Author
Owner

merged.

merged.
steve closed this issue 2026-10-06 13:34:37 -04:00
Sign in to join this conversation.
1 Participants
Notifications
Due Date
No due date set.
Dependencies

No dependencies set.

Reference: steve/skills#30