codex/skills/harness/SKILL.md
Review an agentic system’s configuration and implementation quality. Use when the user wants an opinionated assessment of a system prompt, tool surface, orchestration, guardrails, context handling, or eval setup, and wants concrete recommendations or a redesign plan.
npx skillsauth add tkersey/dotfiles harnessInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Review the agent as an engineering system, not just a prompt.
Harness audits the parts that usually determine whether an agent is actually reliable:
The output should be a judgment with evidence, not a neutral summary.
Use this skill when the user wants to:
Do not use this skill for:
Follow these steps in order.
Start by finding the actual control surface. Look for:
If repository search is needed, run:
./scripts/find_agent_surface.sh
or:
./scripts/find_agent_surface.sh <path>
If the script is unavailable or incomplete, perform an equivalent search manually.
Do not score from one file if the repo clearly has a wider agent surface.
Create a compact evidence map under these headings:
If key evidence is absent, say so before scoring.
Read references/scoring-rubric.md and use it directly.
This rubric is intentionally strict. Do not inflate scores because a demo “basically works.”
Apply hard caps from the rubric when evidence is missing.
Read references/research-basis.md before writing the final verdict.
Your job is not to list generic tradeoffs. Your job is to say what is solid, what is weak, and what should change first.
Examples of acceptable opinions:
Use assets/report-template.md as the report structure.
The report must include:
Judge the prompt on control quality, not eloquence.
Check for:
Penalize:
Review tools as contracts between code and a non-deterministic model.
Check for:
Reward tools that are easy for a new engineer — and for the model — to use correctly.
Penalize tools that dump huge payloads, expose vague parameters, or duplicate one another.
Prefer the simplest architecture that can succeed.
Default stance:
Check for:
Treat guardrails as part of the design, not an afterthought.
Check for:
A mature system should be inspectable.
Check for:
Do not call a system “strong” without runtime evidence.
If the user provides only a prompt, only a schema, or only a partial config:
Never hallucinate maturity.
Recommend the smallest reliability wins first:
Be direct. Be evidence-backed. Do not soften weak design into “just a tradeoff” if it is clearly brittle.
tools
Invokes Apple's macOS 27 fm command-line tool from a local Mac to use the on-device system model or Private Cloud Compute, including instructions, image prompts, schema-constrained JSON, and noninteractive automation. Use when the user asks to run Apple Foundation Models through fm, compare system versus pcc, generate structured output, or automate fm without Swift or an app.
development
Compile historical Codex sessions into governed counterfactual evidence, evaluate an existing owner-applied candidate through blinded paired HCTP trials, and fold observable evidence into RUN, OBSERVE, or STOP. Use for `$hylo`, CRF extraction, counterfactual replay, source-governed direct or historical trials, sealed evidence, paired baseline/candidate evaluation, causal frontiers, or evidence-governed improvement.
testing
Ensure a `ledger` command is available on PATH; materialize, validate, record, replay, and project requested Actuating artifacts without taking semantic or execution authority; coordinate the shared Learnings/Synesthesia/Negative Ledger lifecycle checkpoint and repo-local source-memory reconciliation; address Universalist plans and receipts; and perform pure artifact validation.
testing
Classify and quotient review findings, failing tests, incidents, bug reports, migration failures, and other witnessed falsifiers against accepted intent and the current Construction. Author counterexample-set/v1 without selecting repairs, counting review credit, or granting mutation.