Rewrote the skill so the invoking session only discovers an existing run or launches a fresh planner, then monitors via Herdr read-only commands. Workers keep state in .orchestrate/<slug>/ with run.json and events.jsonl; the tracker holds all handoffs and decisions. Re-invoking the same command recovers monitoring after an orchestrator outage, and an in-progress label plus atomic launch claims prevent duplicate planners.
4.1 KiB
Dispatcher: find or start the planner
Invocation is /skill:orchestrate-herdr <spec-ticket> [worker] (for example, /skill:orchestrate-herdr 34 obelix). The spec issue is required. A named worker is the preferred node for a new planner; if omitted or blank, use the current local worker. Never silently select a different machine.
Find an existing run first
- Confirm the issue exists, is the parent/spec ticket, and is open. If closed, report completion and do not launch anything.
- Resolve the named worker with
herdr machine list; it is the target for a new run. To find an existing run, inspect enabled saved machines using their exact labels/IDs andherdr --machine <machine> agent list. - Read the issue's
in-progresslabel and locate its.orchestrate/<slug>/run.jsonon candidate worker checkouts. Match exact recorded Herdr workspace/pane/agent IDs. Agent listings show runtime identity and status, not the spec issue; use a working directory only to locate a candidate checkout, never as proof of ownership. - If a matching planner is active, monitor it; do not start another. If the matching run exists but the label is missing, verify the run identity, add
in-progress, comment on the reconciliation, and monitor. - If
in-progressis set but no matching agent can be found, do not start a planner. Refresh Herdr and run state to account for propagation delay; if still unmatched, comment on the parent issue with the mismatch and evidence, then stop for human direction. Do not clear the label automatically. - If there is no matching agent and no
in-progresslabel, reconcile any existing run metadata, issue comments, and Git state. If there is no active or ambiguous work, proceed to start a planner. If anything is uncertain, stop rather than risk duplicate work.
For an untracked agent, do not attach to it or assume it belongs to this spec. Flag it for human review if it appears relevant; otherwise leave it untouched.
Start a new planner
- On the selected worker, acquire a per-spec launch claim using an atomic exclusive create under
.orchestrate/. If another invocation holds the claim, reconcile its owner and run state; never steal an uncertain claim. - Create a new sibling pane in the confirmed repository checkout and start a fresh Pi planner agent there. Never use an existing agent pane as the planner. Give the planner the issue number, worker identity, repository path, and references to
run-loop.mdandstate.md. - The planner's first action is to set/verify the spec's
in-progresslabel, before reading tickets or dispatching agents. It then records its Herdr IDs inrun.jsonand begins the run loop. - Confirm the planner reaches
workingand the label is set. Release the launch claim only after confirmed startup. If startup failure is confirmed, release the claim and report the failure. If the outcome is uncertain, retain the claim and stop; do not retry blindly.
Monitor and recover
Use Herdr's read-only commands (agent list, agent wait, agent get, and agent read) to observe agents. Do not use agent attach, send prompts/keys, or resume existing agents. Wait for Herdr state changes and periodically reconcile agent IDs, worker event logs, tracker comments, and Git state.
The same skill invocation is the recovery entry point after an orchestrator outage. Workers continue their issue/PR comments and structured event logging while the orchestrator is unavailable. On re-invocation, inspect current state before doing anything. There is no startup daemon.
If a machine is unreachable, check it with herdr machine status; when authentication is needed, run herdr machine reconnect <label-or-id> and verify connectivity. Retry transient connection failures at most twice after the first attempt. Continue monitoring other reachable work. If the run cannot be safely identified, comment with the blocker and stop; do not dispatch to a different worker.
Herdr machine commands require an enabled saved machine and reachable, API-compatible Herdr server. A failed connection does not prove a mutation failed; reconcile remote state before retrying.