codex/skills/spec-retro/SKILL.md
Historical learning skill for the specification system. Mine multiple prior specs, SGR-v2 receipts, plans, sessions, reports, and churn evidence into concrete updates to `$spec-pipeline` contracts, tools, subagent policy, exemplars, and measurement. Use for `$spec-retro`, improve my spec process from history, analyze spec usage reports, mine plan churn, missing phase impact, repeated gate/challenge/lint failures, or report-to-automation work. Do not use for producing or linting one current spec.
npx skillsauth add tkersey/dotfiles spec-retroInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
$spec-retro improves the specification system from observed practice.
It is separate from $spec-pipeline because it operates across:
Its output is not a current implementation spec.
Its highest-value output is a small set of:
Do not produce more process prose without an actionable rule or measurement.
Use when any are true:
$spec-pipeline runs since the last retro;Do not run after every spec.
Use any available:
spec_governance_receipt / SGR-v2 records;$grill-me snapshots;$spec-pipeline reports;seq results;.ledger/learnings/events.jsonl, previous .ledger/learnings/learnings.jsonl, or legacy .learnings.jsonl during migration;Separate:
retro_evidence:
observed_pattern:
supporting_sources: []
counterevidence: []
confidence:
proposed_automation:
expected_metric_change:
Do not infer a system-wide rule from one ambiguous session.
When given multiple plan/spec files, run:
python codex/skills/spec-retro/scripts/spec_churn_detect.py <files...>
The helper is heuristic. The model still judges whether churn is material.
Emit one compact object:
spec_retro_update:
update_version: SRETRO-v2
report_window: "..."
trigger:
reason: "..."
evidence_refs: []
observed_patterns:
- pattern: "..."
confidence: high | medium | low
counterevidence: []
pipeline_updates:
mode_routing: []
gate_contract: []
challenge_contract: []
fresh_eyes_contract: []
governance_receipt: []
output_templates: []
tool_updates:
churn_detector: []
subagent_policy_updates: []
exemplar_updates: []
measurement_updates:
- metric: "..."
baseline: "..."
expected_change: "..."
next_report_query: "..."
recommendations:
keep: []
revise: []
retire: []
apply_now: yes | no
next_owner: spec-pipeline | spec-retro | tune | seq | user | none
A proposed change is useful only if it can later be graded.
Every recommendation must specify:
hypothesis
expected behavioral change
metric
next report query
keep/revise/retire condition
Preferred historical worker:
spec_retro_miner
Use it only when evidence spans enough artifacts to benefit from a separate historical pass.
The root owns synthesis and recommendation.
tools
Invokes Apple's macOS 27 fm command-line tool from a local Mac to use the on-device system model or Private Cloud Compute, including instructions, image prompts, schema-constrained JSON, and noninteractive automation. Use when the user asks to run Apple Foundation Models through fm, compare system versus pcc, generate structured output, or automate fm without Swift or an app.
development
Compile historical Codex sessions into governed counterfactual evidence, evaluate an existing owner-applied candidate through blinded paired HCTP trials, and fold observable evidence into RUN, OBSERVE, or STOP. Use for `$hylo`, CRF extraction, counterfactual replay, source-governed direct or historical trials, sealed evidence, paired baseline/candidate evaluation, causal frontiers, or evidence-governed improvement.
testing
Ensure a `ledger` command is available on PATH; materialize, validate, record, replay, and project requested Actuating artifacts without taking semantic or execution authority; coordinate the shared Learnings/Synesthesia/Negative Ledger lifecycle checkpoint and repo-local source-memory reconciliation; address Universalist plans and receipts; and perform pure artifact validation.
testing
Classify and quotient review findings, failing tests, incidents, bug reports, migration failures, and other witnessed falsifiers against accepted intent and the current Construction. Author counterexample-set/v1 without selecting repairs, counting review credit, or granting mutation.