skills/standard/shepherd/SKILL.md
Shepherd PRs through CI and reviews to merge readiness. Operates as an iteration loop within the synthesize phase (not a separate HSM phase). Uses assess_stack to check PR health, fix failures, and request approval. Triggers: 'shepherd', 'tend PRs', 'check CI', or /shepherd.
npx skillsauth add lvlup-sw/exarchos shepherdInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
This skill uses VCS operations through Exarchos MCP actions (check_ci, list_prs, merge_pr, get_pr_comments, add_pr_comment, etc.).
These actions automatically detect and route to the correct VCS provider (GitHub, GitLab, Azure DevOps).
No gh/glab/az commands needed — the MCP server handles provider dispatch.
The
merge_prinvoked here is the remote PR merge primitive (synthesize-phase). It is distinct from the integration merge run during the upstreammerge-pendingsubstate — there,serialize_merge(@skills/merge-orchestrator/SKILL.md) is the integration-merge path: it holds a single-writer lease and composes the localgit merge(rawmerge_orchestrateis that composed executor / the non-integration path). This skill never invokesserialize_mergeormerge_orchestrate.
Iterative loop that shepherds published PRs through CI checks and code reviews to merge readiness. Uses the assess_stack composite action for all PR health checks, fixing failures and addressing feedback until the stack is green.
Note: Shepherd is not a separate HSM phase. It operates as a loop within the
synthesizephase. The workflow phase remainssynthesizethroughout the shepherd iteration cycle. Events (shepherd.iteration,ci.status) and theshepherd_statusview track loop progress without requiring a phase transition.
Position in workflow:
synthesize → shepherd (assess → fix → resubmit → loop) → cleanup
^^^^^^^^^ runs within synthesize phase
By default, shepherd seeks to address every piece of feedback on the PR — appending "address all feedback" to the invocation is redundant, because it is the default, not an opt-in. Each iteration drives toward zero outstanding items across all three feedback channels:
CHANGES_REQUESTED review is addressed.assess_stack is advisory: its computeRecommendation only forces fix-and-resubmit on critical/major items, so it can return request-approval while minor comment-reply items remain. Do not treat that as license to skip them — only request approval once every comment-reply action item is addressed (fix or reply). "Address all feedback" is the standing default.
When the pipeline accumulates stale workflows (inactive > 7 days), run @skills/prune/SKILL.md to bulk-cancel abandoned workflows before starting a new shepherd cycle. Discover the stale set with exarchos:exarchos_view({ action: "pipeline", scope: "all" }) — the default view is repo-scoped to the caller's repo and hides legacy rows (workflows started before repo identity was recorded), which are exactly the stale ones this cleanup targets. Safeguards skip workflows with open PRs or recent commits, so active shepherd targets are never touched. A clean pipeline makes shepherd iteration reporting easier to read and reduces noise in the stale-count view.
Activate when:
shepherd or says "shepherd", "tend PRs", "check CI"synthesize completes and PRs are enqueuedsynthesis.prUrl or artifacts.pr)create_pr already ran)Runbook: Each shepherd iteration follows the shepherd-iteration runbook:
exarchos_orchestrate({ action: "runbook", id: "shepherd-iteration" })If runbook unavailable, usedescribeto retrieve action schemas:exarchos_orchestrate({ action: "describe", actions: ["assess_stack"] })
The shepherd loop repeats until all PRs are healthy or escalation criteria are met. Default: 5 iterations.
At the start of each iteration, query quality hints to inform the assessment:
exarchos:exarchos_view({ action: "code_quality", workflowId: "<featureId>" })
regressions is non-empty, include regression context in the status reportconfidenceLevel: 'actionable', surface the suggestedAction in the iteration summarygatePassRate < 0.80 for any skill, flag degrading quality trendsThis step ensures the agent acts on accumulated quality intelligence before polling individual PRs.
Invoke the assess_stack composite action to check all PR dimensions at once:
exarchos:exarchos_orchestrate({
action: "assess_stack",
featureId: "<id>",
prNumbers: [123, 124, 125]
})
The composite action internally handles:
gate.executed events per CI check (feeds CodeQualityView) and ci.status events per PR (feeds ShepherdStatusView). See references/gate-event-emission.md for the event format.Review the returned actionItems and recommendation:
| Recommendation | Action |
|----------------|--------|
| request-approval | Skip to Step 4 only if every comment-reply item is addressed. If any inline comment is still unaddressed, treat as fix-and-resubmit and address it first (see Default Objective). |
| fix-and-resubmit | Proceed to Step 2 |
| wait | Inform user, pause, re-assess after delay |
| escalate | See references/escalation-criteria.md |
Anchor every change to the invariant catalog (evaluation-time Constraints). Before composing any code fix, load .exarchos/invariants.md (entries marked cost-of-load: always-load) and surface a Constraints section naming the invariants the fix must preserve, then probe each proposed change against them. This is the shepherd evaluation-time equivalent of ideate's Phase 0 and uses the same single shared source of truth for the selection rules and devCatalog gating: @skills/ideate/references/constraint-anchoring.md. devCatalog-gated: when .exarchos.yml: invariants.devCatalog: enabled is unset or disabled, surface no Constraints section and fix directly. This makes "when evaluating changes, apply .exarchos/invariants.md" the default — there is no need to request it per invocation.
Before iterating over individual action items, classify them so the loop
knows which to fix inline vs. delegate. Call classify_review_items on
the assessment's actionItems (the comment-reply subset is what the
classifier groups by file; CI-fix and review-address items are passed
through unchanged):
exarchos:exarchos_orchestrate({
action: "classify_review_items",
featureId: "<id>",
actionItems: <actionItems from assess_stack>
})
The result returns groups: ClassificationGroup[] with a recommendation
per group: direct (handle inline), delegate-fixer (spawn the fixer
subagent for batched/HIGH-severity work), or delegate-scaffolder
(cheap subagent for doc nits). Iterate the groups in order, applying
per-group strategy, then consult references/fix-strategies.md for
detailed per-issue-type instructions.
Remediation event protocol (FLYWHEEL):
BEFORE applying a fix, emit remediation.attempted:
exarchos:exarchos_event({
action: "append",
stream: "<featureId>",
event: {
type: "remediation.attempted",
data: { taskId: "<taskId>", skill: "shepherd", gateName: "<failing-gate>", attemptNumber: <N>, strategy: "direct-fix" }
}
})
Apply the fix (CI failure, review comment response, stack restack).
AFTER the next assess confirms the fix resolved the gate, emit remediation.succeeded:
exarchos:exarchos_event({
action: "append",
stream: "<featureId>",
event: {
type: "remediation.succeeded",
data: { taskId: "<taskId>", skill: "shepherd", gateName: "<gate>", totalAttempts: <N>, finalStrategy: "direct-fix" }
}
})
These events feed selfCorrectionRate and avgRemediationAttempts metrics in CodeQualityView.
Action item types:
| Type | Strategy |
|------|----------|
| ci-fix | Read logs, reproduce locally, fix, commit to stack branch |
| comment-reply | Use actionItem.reviewer, normalizedSeverity, file, line, and raw (full original comment) to compose a response. Provider adapters under servers/exarchos-mcp/src/review/providers/ populate the input fields — no manual tier parsing needed. Posting: both PR-level summary comments and per-thread inline replies use the provider-agnostic add_pr_comment orchestrate action — pass a threadId to route a reply into a specific inline thread (via VcsProvider.addReply). Inline thread replies are supported on GitHub; GitLab and Azure DevOps return an explicit "not yet supported" capability signal. |
| review-address | Fix code for CHANGES_REQUESTED, reply to each thread |
| restack | Run git rebase origin/<base>, verify with exarchos_orchestrate({ action: "list_prs" }) |
| escalate | Consult references/escalation-criteria.md |
Every inline review comment must get a reply. The goal is that a human scanning the PR sees every thread has a response.
After fixes are applied, resubmit the stack:
git push --force-with-lease
Re-enable auto-merge if needed:
exarchos_orchestrate({ action: "merge_pr", prId: "<number>", strategy: "squash" })
Return to Step 1 for the next iteration. Track iteration count against the limit (default 5). If the limit is reached without reaching request-approval, escalate per references/escalation-criteria.md.
When assess_stack returns recommendation: 'request-approval' (all checks green, all comments addressed):
Guard: Confirm there are no remaining
comment-replyaction items before requesting approval. Per the Default Objective, every inline comment — including minor/nit-level — must be addressed first. If any remain, return to Step 2 instead of requesting approval.
Invariant-conformance pointer (devCatalog-gated, before merge): When
.exarchos.yml: invariants.devCatalog: enabled, run the review-phase-scopedcheck_invariant_conformanceaction over the PR diff before requesting approval / merge, so the merge-gate read of the architectural invariants matches the diff that is about to land. This is a pointer/affordance, not a hard gate — the action is non-blocking (gate: { blocking: false }); it surfaces conformance findings to fold into the review verdict, it does not stop the merge. It is not a design-time Constraints section (shepherd runs at synthesize/merge, not design-time) — Step 2's Constraints anchoring covers fix composition; this pointer covers the diff-vs-invariants read at the review/merge gate. When the flag is unset ordisabled, skip this step.exarchos:exarchos_orchestrate({ action: "check_invariant_conformance", featureId: "<id>", phase: "review", diff: "<PR diff>" })
Request review via GitHub MCP:
mcp__plugin_github_github__update_pull_request({
owner: "<owner>", repo: "<repo>", pullNumber: <number>,
reviewers: ["<approver>"]
})
Fallback (if MCP token lacks write scope): gh pr edit <number> --add-reviewer <approver>
Report to user:
## Ready for Approval
All CI checks pass. All review comments addressed.
Approval requested from: <approvers>
PRs:
- #123: <url>
- #124: <url>
Run `cleanup` after merge completes.
Track shepherd progress via workflow state:
Initialize:
exarchos:exarchos_workflow({
action: "update",
featureId: "<id>",
updates: {
"shepherd": {
"startedAt": "<ISO8601>",
"currentIteration": 0,
"maxIterations": 5,
"approvalRequested": false
}
}
})
After each iteration: Update currentIteration, record assessment summary and actions taken. On approval: set approvalRequested: true with timestamp and approver list.
For the full transition table, consult @skills/checkpoint/references/phase-transitions.md.
The shepherd skill operates within the synthesize phase and does not drive phase transitions directly.
Use exarchos_workflow({ action: "describe", actions: ["update", "init"] }) for
parameter schemas and exarchos_workflow({ action: "describe", playbook: "feature" })
for phase transitions, guards, and playbook guidance. Use
exarchos_event({ action: "describe", eventTypes: ["shepherd.iteration", "ci.status", "remediation.attempted"] })
for event data schemas before emitting events.
Before emitting any shepherd events, consult references/shepherd-event-schemas.md for full Zod schemas, type constraints, and example payloads. Use exarchos_event({ action: "describe", eventTypes: ["shepherd.iteration", "ci.status"] }) to discover required fields at runtime.
| Event | When | Purpose |
|-------|------|---------|
| shepherd.started | On skill start (emitted by assess_stack) | Audit trail |
| shepherd.iteration | After each assess cycle | Track progress |
| gate.executed | Per CI check (emitted by assess_stack) | CodeQualityView -- gate pass rates |
| ci.status | Per CI check result | ShepherdStatusView -- PR health tracking |
| remediation.attempted | Before applying a fix | selfCorrectionRate metric |
| remediation.succeeded | After fix confirmed | avgRemediationAttempts metric |
| shepherd.approval_requested | On requesting review | Audit trail |
| shepherd.completed | On merge detected (emitted by assess_stack) | Audit trail |
Consult these references for detailed guidance:
references/fix-strategies.md — Fix approaches per issue type, response templates, remediation event emission detailsreferences/escalation-criteria.md — When to stop iterating and escalate to the userreferences/gate-event-emission.md — Event format for gate.executed (now emitted by assess_stack)references/shepherd-event-schemas.md — Full Zod-aligned schemas for all four shepherd lifecycle eventsWhen iteration limits are reached or CI repeatedly fails, consult the escalation runbook:
exarchos_orchestrate({ action: "runbook", id: "shepherd-escalation" })
This runbook provides structured criteria for deciding whether to keep iterating, escalate to the user, or abort the shepherd loop based on iteration count, CI stability, and review status.
| Don't | Do Instead |
|-------|------------|
| Poll CI/reviews directly | Use assess_stack composite action |
| Force-merge with failing CI | Fix the failures first |
| Ignore inline comments | Address every thread with a reply |
| Skip minor/nit comments because assess_stack said request-approval | Address every comment-reply item first — the recommendation is advisory (see Default Objective) |
| Compose fixes without checking invariants | Anchor each change to .exarchos/invariants.md (Step 2, devCatalog-gated) |
| Merge without a diff-vs-invariants read when devCatalog is on | Run check_invariant_conformance over the PR diff before approval/merge (Step 4 pointer, devCatalog-gated, non-blocking) |
| Loop indefinitely | Respect iteration limits, escalate |
| Skip remediation events | Emit remediation.attempted / remediation.succeeded for every fix |
| Push directly to main | All fixes go through stack branches |
After approval is granted and PRs merge, run cleanup to resolve the workflow to completed state.
testing
Create pull request from completed feature branch using GitHub-native stacked PRs. Use when the user says 'create PR', 'submit for review', 'synthesize', or runs /synthesize. Validates branch readiness, creates PR with structured description, and manages merge queue. Do NOT use before review phase completes. Not for draft PRs.
development
Single adversarial review pass over the integrated branch diff — spec-compliance, code quality, and test adequacy judged together by one fresh-context reviewer. Triggers: /review, 'review the changes', or after delegation completes. Emits one verdict (reviews.review.status). Do NOT use for plan-review (that is its own dispatched gate) or for debugging.
testing
Restore and read workflow state after a context break — re-inject workflow phase, task progress, and behavioral guidance into the current session, reconcile state against git reality, and verify whether a workflow exists. Use when the user says 'resume', 'rehydrate', 'where were we', or runs /rehydrate, or when the agent has drifted after context compaction. Do NOT use for saving or mutating state (that is /checkpoint).
testing
Interactively prune stale non-terminal workflows from the pipeline. Use when the user says 'prune workflows', 'clean stale workflows', 'pipeline cleanup', or runs /prune. Runs a dry-run preview, displays candidates with staleness and safeguard skips, prompts the user to proceed/abort/force, then bulk-cancels approved workflows with a workflow.pruned audit event. Safeguards skip workflows with open PRs or recent commits unless force is set.