codex/skills/refine/SKILL.md
Own and apply bounded, evidence-backed optimization of an existing Codex skill. Use after `$tune` supplies STE-v1/SDC-v2 or a complete REFINE-SKILL-v3 brief; inspect the target package, select one smallest intervention, edit only authorized files, preserve stable decision-contract IDs, and retain the named behavioral `$seq` observation query. Not for broad historical diagnosis or system-managed skill optimization.
npx skillsauth add tkersey/dotfiles refineInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
$refine is the user-owned skill optimizer.
$seq gathers evidence
$tune diagnoses the gap and expected decision delta
$refine owns package optimization and editing
Do not delegate optimization to a system-managed skill-optimizer.
$refine may use read-only evidence/modeling subagents supplied by the parent, but root owns all skill-package edits.
Use $refine when the user explicitly authorizes skill changes and one of these exists:
STE-v1 + SDC-v2
REFINE-SKILL-v3 brief
explicit current-turn defect with a complete edit boundary
Do not use $refine to decide whether a skill should change.
Return to $tune when:
Choose exactly one:
inspect
apply
regression
Inspect the target package against an already supplied brief.
Return the smallest viable intervention and outcome-observation plan.
No edits.
Default when explicit edit authority and a complete brief exist.
Inspect, select one intervention, edit, and retain the outcome-observation query.
Repair a previously observed skill failure and install the smallest reproducible guard.
Preferred:
refine_brief:
brief_version: REFINE-SKILL-v3
target_skill:
target_kind:
decision |
execution |
evidence |
orchestration |
mixed
mode:
inspect |
apply |
regression
source_evidence:
packet: STE-v1 | GSD-v2 | user-feedback | observed-behavior
refs: []
gap:
signature:
type:
clause_refs: []
evidence_strength:
recurrence:
expected_delta:
from:
to:
optimization_boundary:
allowed_files: []
forbidden_files: []
protected_contracts: []
intervention_budget:
max_files:
max_new_scripts:
max_new_references:
max_new_contract_clauses:
forbidden_changes: []
smallest_change_hint:
outcome_observation:
query:
evidence_limits: []
Fail closed when the brief does not identify an expected delta and authorized surface.
Read only:
SKILL.md
agents/openai.yaml
references/decision-contract.yaml
brief-authorized references/
brief-authorized scripts/
brief-authorized assets/
Check the worktree before mutation.
Do not mine broad historical sessions inside $refine.
Evaluate only dimensions relevant to the brief:
trigger precision
non-trigger boundary
decision rule
route ownership
mode routing
stop/terminal state
required artifact or receipt
handoff contract
tooling surface
reference/resource placement
metadata/default prompt
decision observability
duplication/deletion
Do not optimize for prose elegance alone.
Select exactly one dominant route:
no_change
trigger_refinement
boundary_refinement
decision_rule_refinement
routing_refinement
workflow_refinement
artifact_refinement
tooling_refinement
resource_refinement
metadata_refinement
observability_refinement
consolidate_or_delete
blocked
A route may touch several files only when they form one coherent intervention, such as:
SKILL.md rule
+ matching decision-contract clause
+ matching agent prompt
Do not combine unrelated improvements into one refinement.
Choose the smallest route that can plausibly produce the expected delta.
Compare candidate interventions in this order:
1. no edit
2. delete or consolidate
3. clarify existing trigger/rule/route
4. repair an existing artifact or operational surface
5. add one narrowly scoped reference
6. add a substantive operation
7. add a new decision-contract clause or receipt
Do not add observability merely because it is possible.
Do not add a script that merely grades skill prose.
When references/decision-contract.yaml exists:
SKILL.md;source_fingerprint after the final package state is known when the local convention supports it.When no contract exists:
SKILL.md under 500 lines.agents/openai.yaml aligned with the final trigger and mission.For a known failure, bind:
observed episode
trigger/clause/route involved
prior bad behavior
expected future behavior
reproduction query
The intervention should address the observed skill failure, not merely change wording.
Examples:
trigger-present but missed activation
prohibited route selected
repeated no-visible-delta ceremony
wrong terminal state
missing required artifact
manual workaround repeated
A decision-oriented skill may emit:
skill_decision_receipt / SDR-v1
Add or require it only when:
An SDR receipt is not proof of a good outcome.
Run the exact named $seq query when supported.
Historical sessions do not retroactively change. Retain the future live query and do not claim that a text edit has already improved behavior.
Every apply/regression run emits:
skill_refinement_receipt:
receipt_version: SRR-v1
target_skill:
target_kind:
source_evidence:
gap_signature:
expected_delta:
from:
to:
selected_route:
alternatives_rejected: []
files_inspected: []
files_changed: []
clauses_changed: []
metadata_disposition:
regenerated |
updated |
verified_unchanged |
not_present
outcome_observation:
current_evidence:
future_query:
residual_uncertainty: []
boundary:
within_boundary: yes | no
expected_delta_addressed: yes | no
The parent may supply:
skill_contract_modeler
skill_decision_provenance_auditor
skill_outcome_skeptic
They provide evidence or skepticism only.
They do not edit files.
$refine root remains the sole writer.
Refined:
- Target:
- Target kind:
- Mode:
- Gap signature:
- Expected delta:
- Selected route:
Changed:
- Files:
- Clauses:
- Metadata:
Outcome observation:
- Current evidence:
- Future query:
Receipt:
- SRR-v1:
Residual uncertainty:
$tune diagnoses; $refine optimizes.skill-optimizer.tools
Invokes Apple's macOS 27 fm command-line tool from a local Mac to use the on-device system model or Private Cloud Compute, including instructions, image prompts, schema-constrained JSON, and noninteractive automation. Use when the user asks to run Apple Foundation Models through fm, compare system versus pcc, generate structured output, or automate fm without Swift or an app.
development
Compile historical Codex sessions into governed counterfactual evidence, evaluate an existing owner-applied candidate through blinded paired HCTP trials, and fold observable evidence into RUN, OBSERVE, or STOP. Use for `$hylo`, CRF extraction, counterfactual replay, source-governed direct or historical trials, sealed evidence, paired baseline/candidate evaluation, causal frontiers, or evidence-governed improvement.
testing
Ensure a `ledger` command is available on PATH; materialize, validate, record, replay, and project requested Actuating artifacts without taking semantic or execution authority; coordinate the shared Learnings/Synesthesia/Negative Ledger lifecycle checkpoint and repo-local source-memory reconciliation; address Universalist plans and receipts; and perform pure artifact validation.
testing
Classify and quotient review findings, failing tests, incidents, bug reports, migration failures, and other witnessed falsifiers against accepted intent and the current Construction. Author counterexample-set/v1 without selecting repairs, counting review credit, or granting mutation.