skills/agentic/context-engineering/SKILL.md
Context window management, token optimization, and memory patterns for efficient multi-agent systems. Use when: optimizing token usage in an agentic pipeline, designing memory scope for short / long-term / episodic state, or applying a context-loading strategy (anticipatory / JIT / hybrid).
npx skillsauth add mikeparcewski/wicked-garden wicked-garden-agentic-context-engineeringInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Techniques for managing context windows, optimizing token usage, and designing efficient memory systems for agentic applications.
Context Window: Maximum tokens an LLM can process in a single request (input + output).
Limits vary by provider and model. Check the active model's documentation for the exact value.
Token Efficiency Matters:
| Pattern | Use when | Pros | Cons | |---------|----------|------|------| | Shared | Agents need synchronized view | Consistency, simple coordination | Contention, single point of failure | | Isolated | Agents operate independently | No contention, parallel execution | Inconsistency possible, harder to coordinate | | Checkpointed | Long-running processes, need recovery | Fault tolerance, replayability | Storage overhead, consistency complexity |
Compress old context into summaries to reduce token usage.
Only load relevant context based on the current task.
Use JSON/structured formats instead of prose to reduce tokens.
Example:
{"name": "John Smith", ...} (compact)Load details only when explicitly needed.
Reference external documents instead of embedding full text.
See refs/selective-loading.md and refs/caching-and-optimization.md for code examples and detailed strategies.
| Memory | Scope | Size | Retention | |--------|-------|------|-----------| | Short-term (working) | Current session/task | 1K-10K tokens | Minutes to hours | | Long-term | Cross-session, permanent | Unbounded (vector DB) | Days to forever | | Episodic | Historical events | Summaries stored | Varies by importance |
See refs/compression-techniques.md for implementation patterns.
Be specific about agent's role and boundaries.
Example:
You are a Python code reviewer specializing in security.
Your job is to identify security vulnerabilities.
You do NOT review style or performance.
Clear, actionable instructions with explicit format.
Bad: "Review this code." Good: "Review for security: 1) SQL injection 2) Input validation 3) Secrets. Output: JSON with vulnerabilities."
Specify exact output format to reduce tokens.
Show examples for complex tasks.
See refs/selective-loading.md for detailed prompting patterns.
| Strategy | Pros | Cons | |----------|------|------| | Anticipatory | Faster response time (load before needed) | May load unnecessary data | | Just-in-Time (JIT) | Minimal token usage (load only when needed) | Latency on each request | | Hybrid | Balanced (core context + JIT for task-specific) | More complex implementation |
Track input and output tokens separately. Rates vary by model (typically $0.003-0.075 per 1K tokens).
Set hard token limits per agent/session to prevent runaway costs.
Track costs per agent to identify expensive components.
See refs/cost-calculation-budget.md and refs/cost-optimization-reporting.md for detailed cost strategies.
Sequential Pattern: Pass only output of previous agent, not entire chain.
Hierarchical Pattern: Parent gets summaries from children, children get only relevant task context.
Collaborative Pattern: Shared context (compressed), each agent adds only delta.
Autonomous Pattern: Minimal shared context, isolated context per agent.
refs/compression-techniques.md - Conversation summarization, deduplication, entity compressionrefs/selective-loading.md - Relevance filtering, time decay, token-budgeted retrievalrefs/caching-and-optimization.md - Prompt caching, semantic caching, batching, cost-aware model selectionrefs/cost-calculation-budget.md - Token pricing, cost calculation, budget managementrefs/cost-optimization-reporting.md - Cost estimation, optimization strategies, reportingdevelopment
Pattern-conformance agent-half: evaluates a produced artifact or diff against a set of architectural/design pattern rules from the conformance-rule store (wicked_governance schema). Returns structured findings with rule ID, severity, and rationale — the deterministic half (mechanical rule recall) is done by the guard pipeline; this is the semantic evaluation step. Triggered by: the guard_pipeline `outgov_pattern` check (session-close), or explicitly by an engineering review when WICKED_OUTGOV_RULES_DIR is populated. NOT a replacement for the full `engineering` review skill — focuses only on conformance to stored Pattern rules; architecture and code-quality checks live in the `engineering` skill. Semantic evaluation reuses `wicked-garden-qe-semantic-reviewer` as the designated agent-half evaluator (per garden#983 spec). This skill is the orchestrating wrapper that loads applicable Pattern rules and delegates the per-rule semantic judgment to qe-semantic-reviewer.
tools
The FOUNDATIONAL domain-model capability: extract a codebase's domain — testable business rules (with confidence + provenance), entities, requirements — as a schema-conformant model on the estate graph. The workers annotate the store; wicked-core reads it and builds the requirements graph, coverage-gating fail-closed. Steers three fork workers. A shared substrate, not a modernization tool. The `modernize` archetype DERIVES from it; build / migrate / review / specify / explore consume the SAME domain model — none OWN it. Understanding a codebase's domain is upstream of almost everything else garden does. Use when: "extract the business rules / domain model from this codebase", "build a requirements graph from the code", "what does this system actually require", "reverse-engineer the domain before we build/port/migrate". Works on ANY codebase (modern or legacy) — the value is the domain model, not the porting. NOT the code transform itself (that is the archetype consuming this model). This skill produces the DOMAIN MODEL, not new code.
development
Domain-graph fork worker for the modernize archetype. Groups the estate's Louvain communities into business domains, attaches each requirement to its cluster (advisory cluster_id provenance), and invokes wicked-core's domain-graph build (which reads the annotated estate store, recomputes coverage fail-closed, and builds the requirements graph) — then validates core's output against the vendored schema. Use when: dispatched by wicked-garden-domain after rule extraction to turn a flat rule set into cluster-keyed domains; "group these into domains", "build the requirements graph", "translate clusters into a domain model". NOT for mining the rules themselves (that is domain-extractor) or threat-modeling (that is domain-coverage).
tools
Rule-extraction fork worker for the FOUNDATIONAL domain-model capability. Mines testable business rules from a codebase — each with a numeric confidence and a provenance{source, ref, source_kinds} — and annotates them into the estate store so wicked-core can build the domain-model requirements graph (coverage-gated). This is a substrate, not a modernization tool: the `modernize` archetype DERIVES from it, and build / migrate / review / specify / explore can consume the same domain model — none OWN it. Use when: dispatched by wicked-garden-domain to mine the business_rules of a codebase (or a module); "extract the domain rules", "what does this system require", building the requirements half of a domain model. NOT for grouping into domains (that is domain-modeler) or judging coverage (that is domain-coverage — a seat-distinct evaluator).