skills/data/SKILL.md
Data engineering, analysis, and ML toolkit. One skill, routed sub-actions: analyze (interactive EDA on a CSV/Excel/data file), profile / validate / quality (schema-level data engineering ops), ml review / ml pipeline (model review + training-pipeline design), pipeline design / pipeline review (ETL architecture), and ontology (recommend Schema.org, Dublin Core, DCAT, FOAF, GoodRelations, SKOS, or a custom shape for a dataset). Use when: "analyze this CSV/Excel/data file" / "explore this dataset" / one-off data exploration; "profile this dataset" / "validate data against a schema" / "generate a data quality report" (completeness, uniqueness, validity); "review this ML model" / "design an ML training pipeline"; "design a data pipeline" / "review this ETL" / "optimize data processing"; "recommend an ontology for this data" / "map columns to a public ontology". Replaces the former /wicked-garden:data:* commands (analyze, data, ml, ontology, pipeline). NOT for delegated subagent work — dispatch the wicked-garden-data-engineer fork skill for that.
npx skillsauth add mikeparcewski/wicked-garden wicked-garden-dataInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Every sub-action runs inline (no dispatch): parse the arguments, pre-read the data where noted, load the Tier-3 rubric from refs/, apply it, and emit structured markdown with tables and prioritized findings. For delegated or parallel worker execution, dispatch the wicked-garden-data-engineer fork skill instead.
| Sub-action | Use for | Rubric |
|------------|---------|--------|
| analyze <file-path> [--focus stats\|quality\|warehouse\|ml] [--context <file>] [--refresh] [--scenarios] | one-off exploration of a CSV/Excel/data file | refs/analyze.md |
| profile <path> | dataset structure + quality profile | refs/data.md |
| validate --schema <schema> --data <path> | schema validation | refs/data.md |
| quality <path> | quality report (completeness, uniqueness, validity) | refs/data.md |
| ml review <path> | ML model review | refs/ml.md |
| ml pipeline --type <classification\|regression\|ranking> | training-pipeline design | refs/ml.md |
| pipeline design --source <src> --target <tgt> [--frequency <freq>] | ETL pipeline design | refs/pipeline.md |
| pipeline review <path> | ETL pipeline review | refs/pipeline.md |
| ontology <file-path> | ontology recommendation for a dataset | inline (§ Ontology) |
Detailed templates and examples: refs/analysis-templates.md, refs/ml-templates.md, refs/pipeline-templates.md, refs/data-examples.md.
Interactive analysis on a CSV/Excel/data file. Use for one-off data
exploration. NOT for schema-level checks (use profile / validate /
quality) or pipeline review (use pipeline review).
<file-path>, --focus (stats|quality|warehouse|ml, default
stats), --context, --refresh, --scenarios.Read("${CLAUDE_PLUGIN_ROOT}/skills/data/refs/analyze.md") — the EDA
rubric, quality/warehouse/ml modes, insight pattern, and output format.--focus mode and emit the analysis.Schema-level engineering ops on a dataset. NOT for interactive exploration
(use analyze) or ML pipeline review (use ml).
<path>,
--schema for validate).Read("${CLAUDE_PLUGIN_ROOT}/skills/data/refs/data.md") — the profile,
validate, and quality rubrics with output formats and quality thresholds.Optional scripted paths (deterministic profiling/validation):
sh "${CLAUDE_PLUGIN_ROOT}/scripts/_python.sh" "${CLAUDE_PLUGIN_ROOT}/scripts/data/data_profiler.py" \
--input data.csv --output profile.json
sh "${CLAUDE_PLUGIN_ROOT}/scripts/_python.sh" "${CLAUDE_PLUGIN_ROOT}/scripts/data/schema_validator.py" \
--schema schemas/expected.json \
--data data/actual.csv
For files >1GB, use the analyze sub-action for efficient SQL-based profiling
via DuckDB.
ML model review and training-pipeline design. NOT for ETL pipeline design
(use pipeline) or data profiling (use profile).
<path> for review,
--type for pipeline).review, read model files at <path>.Read("${CLAUDE_PLUGIN_ROOT}/skills/data/refs/ml.md") — the model review
checklist, pipeline design template, deployment readiness checklist, and
MLOps standards.Data pipeline design and review. NOT for ML training pipelines (use
ml pipeline) or one-off file analysis (use analyze).
review, read pipeline
files at <path>. For design, capture --source, --target,
--frequency.Read("${CLAUDE_PLUGIN_ROOT}/skills/data/refs/pipeline.md") — the design
checklist, review rubric with P1/P2/P3 findings, pattern selection, and
engineering standards.Sample a dataset (CSV/Excel/Parquet/JSON) and recommend matching public
ontologies (Schema.org, Dublin Core, DCAT, FOAF, GoodRelations, SKOS) or a
custom shape. Use this for ontology mapping. NOT for interactive analysis
(use analyze) or quality reports (use quality).
Arg parse — extract file-path from the arguments.
Run recommender:
cd "${CLAUDE_PLUGIN_ROOT}" && uv run python scripts/_run.py scripts/data/ontology_recommender.py "${file_path}"
Present the script's match table, column-mapping suggestions, and any custom-ontology fallback inline.
development
Pattern-conformance agent-half: evaluates a produced artifact or diff against a set of architectural/design pattern rules from the conformance-rule store (wicked_governance schema). Returns structured findings with rule ID, severity, and rationale — the deterministic half (mechanical rule recall) is done by the guard pipeline; this is the semantic evaluation step. Triggered by: the guard_pipeline `outgov_pattern` check (session-close), or explicitly by an engineering review when WICKED_OUTGOV_RULES_DIR is populated. NOT a replacement for the full `engineering` review skill — focuses only on conformance to stored Pattern rules; architecture and code-quality checks live in the `engineering` skill. Semantic evaluation reuses `wicked-garden-qe-semantic-reviewer` as the designated agent-half evaluator (per garden#983 spec). This skill is the orchestrating wrapper that loads applicable Pattern rules and delegates the per-rule semantic judgment to qe-semantic-reviewer.
tools
The FOUNDATIONAL domain-model capability: extract a codebase's domain — testable business rules (with confidence + provenance), entities, requirements — as a schema-conformant model on the estate graph. The workers annotate the store; wicked-core reads it and builds the requirements graph, coverage-gating fail-closed. Steers three fork workers. A shared substrate, not a modernization tool. The `modernize` archetype DERIVES from it; build / migrate / review / specify / explore consume the SAME domain model — none OWN it. Understanding a codebase's domain is upstream of almost everything else garden does. Use when: "extract the business rules / domain model from this codebase", "build a requirements graph from the code", "what does this system actually require", "reverse-engineer the domain before we build/port/migrate". Works on ANY codebase (modern or legacy) — the value is the domain model, not the porting. NOT the code transform itself (that is the archetype consuming this model). This skill produces the DOMAIN MODEL, not new code.
development
Domain-graph fork worker for the modernize archetype. Groups the estate's Louvain communities into business domains, attaches each requirement to its cluster (advisory cluster_id provenance), and invokes wicked-core's domain-graph build (which reads the annotated estate store, recomputes coverage fail-closed, and builds the requirements graph) — then validates core's output against the vendored schema. Use when: dispatched by wicked-garden-domain after rule extraction to turn a flat rule set into cluster-keyed domains; "group these into domains", "build the requirements graph", "translate clusters into a domain model". NOT for mining the rules themselves (that is domain-extractor) or threat-modeling (that is domain-coverage).
tools
Rule-extraction fork worker for the FOUNDATIONAL domain-model capability. Mines testable business rules from a codebase — each with a numeric confidence and a provenance{source, ref, source_kinds} — and annotates them into the estate store so wicked-core can build the domain-model requirements graph (coverage-gated). This is a substrate, not a modernization tool: the `modernize` archetype DERIVES from it, and build / migrate / review / specify / explore can consume the same domain model — none OWN it. Use when: dispatched by wicked-garden-domain to mine the business_rules of a codebase (or a module); "extract the domain rules", "what does this system require", building the requirements half of a domain model. NOT for grouping into domains (that is domain-modeler) or judging coverage (that is domain-coverage — a seat-distinct evaluator).