skills/platform/health/SKILL.md
System health overview from discovered observability sources. Aggregates errors, performance metrics, and SLO status across services. Correlates with deployments and code changes. Use for proactive health monitoring and post-deployment validation. Use when: checking aggregated system health, validating a post-deployment state, correlating production status with recent changes, "how is production", "check system health", or any former /wicked-garden:platform:health invocation. NOT for plugin-level diagnostics (use the observability sub-skill) or distributed tracing (use the platform domain skill's traces action).
npx skillsauth add mikeparcewski/wicked-garden wicked-garden-platform-healthInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Aggregate system health from discovered observability sources with deployment correlation.
NOT for plugin-level diagnostics (use the observability sub-skill) or
distributed tracing (use the traces action of the platform domain skill).
Invoked as [service name or 'all']:
all for full assessment.Read("${CLAUDE_PLUGIN_ROOT}/skills/platform/health/refs/health.md") —
discovery checklist, per-source assessment steps, fallback code-analysis
approach, common patterns, and output format. See refs/sources.md for
detailed capability discovery patterns.ListMcpResourcesTool, query each, correlate with recent deployments, and
produce the health report with severity classification
(HEALTHY / DEGRADED / CRITICAL) and prioritized recommendations.Use capability-based discovery to find available integrations:
# Discover available integrations via capability detection
# Scan for capabilities by analyzing server descriptions and resources:
# - error-tracking capability: Exception/error tracking and reporting
# - apm capability: Application performance monitoring and metrics
# - logging capability: Log aggregation, search, and analysis
# - tracing capability: Distributed tracing and service mapping
# - telemetry capability: Metrics collection and custom instrumentation
For each discovered source, collect:
HEALTHY: All metrics within SLO, no active alerts, stable trends DEGRADED: Some metrics elevated, minor alerts, or negative trends CRITICAL: SLO violations, critical alerts, or severe degradation
Check for recent changes that might impact health:
Based on health status:
This skill discovers integrations at runtime based on capability:
| Capability | What to Look For | Provides | |------------|------------------|----------| | error-tracking | Exception tracking, error reporting, crash analytics | Error rates, stack traces, user impact | | apm | Performance monitoring, service metrics, observability | Latency, throughput, service health | | logging | Log aggregation, log search, log analysis | Log aggregation, search, patterns | | tracing | Distributed tracing, request tracing, trace analysis | Distributed traces, dependencies | | telemetry | Metrics collection, custom instrumentation, time-series data | Custom metrics, instrumentation |
Fallback: If no integrations found, perform local analysis via wicked-garden:search for error patterns in code.
See refs/sources.md for detailed capability discovery patterns.
## System Health Report
**Overall Status**: [HEALTHY | DEGRADED | CRITICAL]
**Assessment Time**: {timestamp}
**Data Sources**: {list of integrations used}
### Health Summary
| Service | Status | Error Rate | Latency (p95) | SLO Status |
|---------|--------|------------|---------------|------------|
| {service} | {status} | {rate} | {latency} | {✓ or ✗} |
### Issues Detected
[For each issue]
**{Service}: {Issue Description}**
- Severity: [CRITICAL | HIGH | MEDIUM | LOW]
- Started: {timestamp}
- Metric: {specific metric and values}
- Pattern: {error pattern or behavior}
- Correlation: {deployment or change if found}
- Blast Radius: {impact scope}
### Trends (24h)
- Error Rates: {trend with percentage}
- Latency: {trend with percentage}
- Traffic: {trend with percentage}
### Recommendations
**Immediate**:
{critical actions needed now}
**Short-term**:
{optimizations and improvements}
**Capacity**:
{capacity planning insights}
Error rates or latency increase after deployment. Correlate metrics with deployment time and consider rollback.
Metrics slowly degrading over hours/days. Investigate memory leaks, growing data, cache efficiency.
Performance degrades with traffic spikes. Check capacity utilization and scaling policies.
Single service failure causes downstream issues. Use traces to identify root cause and implement circuit breakers.
When crew enters build phase:
Emit events:
observe:health:checked:successobserve:health:degraded:warningobserve:health:critical:failureWhen debugging issues, provide observability context:
development
Pattern-conformance agent-half: evaluates a produced artifact or diff against a set of architectural/design pattern rules from the conformance-rule store (wicked_governance schema). Returns structured findings with rule ID, severity, and rationale — the deterministic half (mechanical rule recall) is done by the guard pipeline; this is the semantic evaluation step. Triggered by: the guard_pipeline `outgov_pattern` check (session-close), or explicitly by an engineering review when WICKED_OUTGOV_RULES_DIR is populated. NOT a replacement for the full `engineering` review skill — focuses only on conformance to stored Pattern rules; architecture and code-quality checks live in the `engineering` skill. Semantic evaluation reuses `wicked-garden-qe-semantic-reviewer` as the designated agent-half evaluator (per garden#983 spec). This skill is the orchestrating wrapper that loads applicable Pattern rules and delegates the per-rule semantic judgment to qe-semantic-reviewer.
tools
The FOUNDATIONAL domain-model capability: extract a codebase's domain — testable business rules (with confidence + provenance), entities, requirements — as a schema-conformant model on the estate graph. The workers annotate the store; wicked-core reads it and builds the requirements graph, coverage-gating fail-closed. Steers three fork workers. A shared substrate, not a modernization tool. The `modernize` archetype DERIVES from it; build / migrate / review / specify / explore consume the SAME domain model — none OWN it. Understanding a codebase's domain is upstream of almost everything else garden does. Use when: "extract the business rules / domain model from this codebase", "build a requirements graph from the code", "what does this system actually require", "reverse-engineer the domain before we build/port/migrate". Works on ANY codebase (modern or legacy) — the value is the domain model, not the porting. NOT the code transform itself (that is the archetype consuming this model). This skill produces the DOMAIN MODEL, not new code.
development
Domain-graph fork worker for the modernize archetype. Groups the estate's Louvain communities into business domains, attaches each requirement to its cluster (advisory cluster_id provenance), and invokes wicked-core's domain-graph build (which reads the annotated estate store, recomputes coverage fail-closed, and builds the requirements graph) — then validates core's output against the vendored schema. Use when: dispatched by wicked-garden-domain after rule extraction to turn a flat rule set into cluster-keyed domains; "group these into domains", "build the requirements graph", "translate clusters into a domain model". NOT for mining the rules themselves (that is domain-extractor) or threat-modeling (that is domain-coverage).
tools
Rule-extraction fork worker for the FOUNDATIONAL domain-model capability. Mines testable business rules from a codebase — each with a numeric confidence and a provenance{source, ref, source_kinds} — and annotates them into the estate store so wicked-core can build the domain-model requirements graph (coverage-gated). This is a substrate, not a modernization tool: the `modernize` archetype DERIVES from it, and build / migrate / review / specify / explore can consume the same domain model — none OWN it. Use when: dispatched by wicked-garden-domain to mine the business_rules of a codebase (or a module); "extract the domain rules", "what does this system require", building the requirements half of a domain model. NOT for grouping into domains (that is domain-modeler) or judging coverage (that is domain-coverage — a seat-distinct evaluator).