skills/crew/evidence-validation/SKILL.md
Validates that completed task descriptions include required evidence fields at the appropriate level for the task's complexity score. Three tiers (low, medium, high) map to complexity ranges 1-2, 3-4, 5-7. Use when: validating a TaskUpdate description before marking complete, or checking that evidence fields match the task's complexity tier.
npx skillsauth add mikeparcewski/wicked-garden evidence-validationInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Check that a task description has required evidence before marking complete.
{
"valid": true|false,
"missing": ["Human-readable label for missing field", ...],
"warnings": ["Advisory warning (non-blocking)", ...]
}
| Complexity | Tier | Required Fields | |------------|------|-----------------| | 0-2 | low | test_results + code_diff | | 3-4 | medium | test_results + code_diff + verification | | 5-7 | high | test_results + code_diff + verification + performance + assumptions |
Tiers are cumulative: high = medium + performance + assumptions.
missingwarnings for advisory itemsvalid = true only when missing is emptyFor any complexity >= 3, if no assumptions are detected, add this warning
(non-blocking, does not affect valid):
"Consider documenting assumptions (## Assumptions section) for medium/high complexity tasks to help reviewers."
This warning is suppressed if assumptions is already in missing (i.e.,
high-tier tasks where it is required).
Each field is detected by scanning the task description for natural language indicators. Detection is text analysis — look for the patterns described in Evidence Schema. A single indicator match is sufficient to mark a field as present.
Detection is case-insensitive.
| Field | Label used in missing |
|-------|------------------------|
| test_results | "Test results (e.g. '- Test: test_name — PASS/FAIL')" |
| code_diff | "Code diff reference (e.g. '- Code diff: ...' or '- File: path — modified/created')" |
| verification | "Verification step (e.g. '- Verification: curl ... returns 200' or command output)" |
| performance | "Performance data (e.g. latency, throughput, benchmark results)" |
| assumptions | "Documented assumptions (e.g. '## Assumptions' section or '- Assumption:')" |
Well-structured task descriptions follow this format:
{original task description}
## Outcome
{what was accomplished}
## Evidence
- Test: {test name} — PASS/FAIL
- File: {path} — modified/created/deleted
- Verification: {command + output}
- Performance: {metric} (complexity >= 5 only)
- Benchmark: {tool + result} (complexity >= 5 only)
## Assumptions
- {assumption 1}
See Evidence Schema for full detection patterns and Test Evidence Requirements for QE artifact requirements per test type.
When a task description includes an ## Acceptance Criteria section (AC-4.5), validation must confirm that each listed AC has a corresponding evidence entry — not just that evidence fields are present anywhere in the description.
## Acceptance Criteria section. Each - [ ] AC-N: or - AC-N: line is a criterion that requires evidence.## Evidence section. Each - AC-N: line is a candidate match.AC-1 in criteria → AC-1: in evidence). Label matching is case-insensitive.If any AC is missing a matching evidence entry, add to missing:
"AC-N evidence: {criterion text} — no matching evidence entry found"
If an evidence entry exists but contains only an assertion without artifact reference, add to warnings:
"AC-N evidence is vague: '{entry text}' — provide file path, test output, or diff"
AC-evidence mapping is performed in addition to the tier-based field checks. A task can pass tier validation (evidence fields present) but still fail AC-evidence mapping if individual criteria lack coverage.
development
Pattern-conformance agent-half: evaluates a produced artifact or diff against a set of architectural/design pattern rules from the conformance-rule store (wicked_governance schema). Returns structured findings with rule ID, severity, and rationale — the deterministic half (mechanical rule recall) is done by the guard pipeline; this is the semantic evaluation step. Triggered by: the guard_pipeline `outgov_pattern` check (session-close), or explicitly by an engineering review when WICKED_OUTGOV_RULES_DIR is populated. NOT a replacement for the full `engineering` review skill — focuses only on conformance to stored Pattern rules; architecture and code-quality checks live in the `engineering` skill. Semantic evaluation reuses `wicked-garden-qe-semantic-reviewer` as the designated agent-half evaluator (per garden#983 spec). This skill is the orchestrating wrapper that loads applicable Pattern rules and delegates the per-rule semantic judgment to qe-semantic-reviewer.
tools
The FOUNDATIONAL domain-model capability: extract a codebase's domain — testable business rules (with confidence + provenance), entities, requirements — as a schema-conformant model on the estate graph. The workers annotate the store; wicked-core reads it and builds the requirements graph, coverage-gating fail-closed. Steers three fork workers. A shared substrate, not a modernization tool. The `modernize` archetype DERIVES from it; build / migrate / review / specify / explore consume the SAME domain model — none OWN it. Understanding a codebase's domain is upstream of almost everything else garden does. Use when: "extract the business rules / domain model from this codebase", "build a requirements graph from the code", "what does this system actually require", "reverse-engineer the domain before we build/port/migrate". Works on ANY codebase (modern or legacy) — the value is the domain model, not the porting. NOT the code transform itself (that is the archetype consuming this model). This skill produces the DOMAIN MODEL, not new code.
development
Domain-graph fork worker for the modernize archetype. Groups the estate's Louvain communities into business domains, attaches each requirement to its cluster (advisory cluster_id provenance), and invokes wicked-core's domain-graph build (which reads the annotated estate store, recomputes coverage fail-closed, and builds the requirements graph) — then validates core's output against the vendored schema. Use when: dispatched by wicked-garden-domain after rule extraction to turn a flat rule set into cluster-keyed domains; "group these into domains", "build the requirements graph", "translate clusters into a domain model". NOT for mining the rules themselves (that is domain-extractor) or threat-modeling (that is domain-coverage).
tools
Rule-extraction fork worker for the FOUNDATIONAL domain-model capability. Mines testable business rules from a codebase — each with a numeric confidence and a provenance{source, ref, source_kinds} — and annotates them into the estate store so wicked-core can build the domain-model requirements graph (coverage-gated). This is a substrate, not a modernization tool: the `modernize` archetype DERIVES from it, and build / migrate / review / specify / explore can consume the same domain model — none OWN it. Use when: dispatched by wicked-garden-domain to mine the business_rules of a codebase (or a module); "extract the domain rules", "what does this system require", building the requirements half of a domain model. NOT for grouping into domains (that is domain-modeler) or judging coverage (that is domain-coverage — a seat-distinct evaluator).