aops-pkb/skills/verify/SKILL.md
Judgement-based QA pass. Does this artifact meet its goal and serve its user? Demands excellence, not compliance. Owned by marsha; reads the spec's Fitness Rubric (designed upstream via /design-rubric).
npx skillsauth add nicsuzor/academicops verifyInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Conduct rigorous QA reviews of artifacts to ensure correctness, complete implementation, and fitness for purpose.
Personality binding — earmarking. This skill is earmarked to marsha: the "assume it's broken" default posture below (Core Directives) is a direct expression of her broken-until-proven-otherwise judgment register (specs/agents/marsha.md), not a generic QA checklist any agent could apply as well. Another agent could technically execute the steps, but the disposition the steps depend on — refusing to credit an artifact until it's independently proven, holding the line against agents' default eagerness to declare victory — is marsha's, and running this skill under a different disposition would silently soften the bar it's designed to hold.
Before you read a single line of the diff, judge the premise from the task + diffstat alone and write the sharp principal's one-sentence snap reaction — "was this a good idea, in this shape?" — verbatim, as forcing-check item 0. You cannot emit a PASS verdict without it; a bad premise is a FAIL regardless of test coverage (green tests are the expected surface of a bad premise, not a mitigant). Diffstat-first ordering is mandatory — reading the code first is exactly what lets a clean, well-tested surface launder a bad premise.
Full definition, the verbatim prompt, the never-a-checklist hard rule, and the worked specimen live in the canonical reference: [[premise-test.md]]. (FAIL is the local rejection token here; the arch-fit lens emits 🔴 REJECT for the same call.)
Default posture: assume it's broken. The burden is on the artifact to prove it works — not on you to prove it doesn't.
PASS, FAIL, or REVISE.## Fitness Rubric. (If missing on a fitness task, return REVISE — fitness rubric missing)..agents/rules/RULES.md exists in this repo, read it before judging. Apply its rules with the same class/instance discipline as AXIOMS.md. Project-rule violations belong under Process Compliance in the report, cited by {#slug}. RULES.md is not the only standard: for a content/instruction artifact (skill, agent body, prompt, doc, spec) also identify the skill that owns its quality standard for that artifact type and verify against it — e.g. /craft for instruction / agent-definition / skill / prompt edits. The governing standard often lives in a skill, not RULES.md.FAIL regardless of test coverage; you cannot reach PASS without writing it.DERIVER_MISSING, N/A, TODO). Fail if primary value-signals are missing./design-rubric self-instance requirement).For any artifact with computed, aggregated, or derived output (dashboards, reports, metrics), trace source → output: confirm the source is real, populated, and fresh; independently cross-verify the values against that source; disable any fallback to prove the primary path works alone (a fallback silently masks a broken primary); and check behaviour under load. The question is not "did output appear?" but "is this the RIGHT data?" — plausible-looking output is the most dangerous kind of incorrect output.
Stop evaluation immediately and write a FAIL verdict if any of the following occur:
FAIL regardless of green tests; test-passing is the expected surface of this failure, not a mitigant.{variable}, TODO, FIXME) in production.try/except without logging).Output reports exactly in this format:
## Verification Report
**Bar:** [mechanical / fitness / mixed]
**Verdict:** [PASS / FAIL / REVISE]
### Concrete observations
[Observed bugs/defects, file paths, line numbers, and log excerpts]
### Forcing checks
0. **Premise test (before reading the diff):** [verbatim sharp-principal reaction from task + diffstat alone — "was this a good idea, in this shape?" A bad premise -> FAIL regardless of tests; cannot reach PASS without this line]
1. **Sentinel/empty-state audit:** [count + list of sentinels/placeholders. If primary signals absent -> FAIL]
2. **Principal's-eye top-line read:** [headline element quoted, and whether correct]
3. **Floor vs ceiling:** [verbatim "exceptional, or merely working?"]
### Process compliance
[Project-rule violations cited by `{#slug}` from `.agents/rules/RULES.md` if present, or "RULES.md absent — skipped"]
### Judgement
[Prose evaluation against AC, Red Flags, and/or Fitness Rubric dimensions]
### Recommendation
[If FAIL/REVISE: specific remediation steps and user impact]
For web applications:
$AOPS_SESSIONS/qa-screenshots/YYYY-MM-DD/.data-ai
Canonical session close — commit, push, PR, release_task, reflection blocks, handover. Use /dump for emergency bail (no commit/PR/reflection).
data-ai
Emergency session bail — fast resume task + short handover, no commit/PR/reflection. For when you (or the user) need a clean context now. Use /end-session for canonical close.
data-ai
Daily note lifecycle — compose and maintain a factual daily note. Reports the state of the day; does not prioritise or recommend. SSoT for daily note structure.
testing
Launder supervisor/worker task-log output into a Nic-facing narrative — what happened, where things are headed, and what (if anything) is genuinely his to decide. Never relays raw process detail (worker IDs, thread pointers, log paths) or verbatim task-log stream-of-consciousness.