codex/skills/guards/SKILL.md
Generate provider-agnostic AI agent guardrail blueprints and control matrices from a use case. Use when designing or reviewing agent safety architecture, prompt-injection and tool-misuse defenses, risk-tiered human approval gates, or auditable enterprise guardrail policies using industry patterns across top providers.
npx skillsauth add tkersey/dotfiles guardsInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Turn ambiguous "add guardrails" requests into an implementation-ready guardrail architecture. Produce provider-agnostic outputs that application engineers can execute without choosing a model vendor first.
Always return a structured plan plus a control matrix. The response must include these headings in this order:
Never output provider-specific SDK instructions in default mode.
If a user explicitly requests provider mappings, keep them in a separate appendix titled Provider Mapping (Optional).
Model requests as GuardrailBlueprintRequest v1.
Required fields:
use_case: one-sentence workload descriptionactor_profile: who can trigger the agent and from wheretool_surface: tools or side-effecting actions the agent can invokedata_sensitivity: public | internal | confidential | regulatedautonomy_level: assistant | semi_autonomous | autonomousregulatory_context: baseline policy contextrisk_tolerance: low | medium | highDefaults:
regulatory_context = US enterprise baselinerisk_tolerance = mediumfreshness_mode = pattern_onlyWhen required fields are missing, ask only the minimum judgment-call questions needed to classify risk tier and tool boundary.
Map every threat scenario across this layer order:
Each scenario must include at least one control in every applicable layer.
Apply this default policy unless the user overrides it:
low: monitored allow with explicit logging and anomaly detectionmedium: allow with guardrails and sampled human reviewhigh: fail-closed for side-effecting actions unless explicit human approvalcritical: fail-closed by default with mandatory human approval and dual-control where possibleHard rule:
GuardrailBlueprintRequest v1.references/industry_patterns.md.references/blueprint_template.md.Use one mode and state it in the output:
pattern_only (default): synthesize provider-agnostic industry patterns.hybrid: cite sources for high-risk controls and decision-critical claims.strict_source_cited: cite sources for all major control recommendations.When the prompt asks for "latest" provider specifics, verify against primary docs before finalizing.
Before completing, verify all gates:
pattern_only, hybrid, or strict_source_cited).references/blueprint_template.mdreferences/industry_patterns.mdtools
Invokes Apple's macOS 27 fm command-line tool from a local Mac to use the on-device system model or Private Cloud Compute, including instructions, image prompts, schema-constrained JSON, and noninteractive automation. Use when the user asks to run Apple Foundation Models through fm, compare system versus pcc, generate structured output, or automate fm without Swift or an app.
development
Compile historical Codex sessions into governed counterfactual evidence, evaluate an existing owner-applied candidate through blinded paired HCTP trials, and fold observable evidence into RUN, OBSERVE, or STOP. Use for `$hylo`, CRF extraction, counterfactual replay, source-governed direct or historical trials, sealed evidence, paired baseline/candidate evaluation, causal frontiers, or evidence-governed improvement.
testing
Ensure a `ledger` command is available on PATH; materialize, validate, record, replay, and project requested Actuating artifacts without taking semantic or execution authority; coordinate the shared Learnings/Synesthesia/Negative Ledger lifecycle checkpoint and repo-local source-memory reconciliation; address Universalist plans and receipts; and perform pure artifact validation.
testing
Classify and quotient review findings, failing tests, incidents, bug reports, migration failures, and other witnessed falsifiers against accepted intent and the current Construction. Author counterexample-set/v1 without selecting repairs, counting review credit, or granting mutation.