plugins/pm-ai/skills/agent-design-review/SKILL.md
Review an LLM agent design and find where it will be unreliable, expensive, or unsafe. Use when asked to review an agent architecture, critique a multi-step/tool-using agent, debug an agent that loops or goes off-task, or harden an agent before launch. Produces a structured review — task fit, control flow, tools, memory/context, failure handling, cost, and safety — with prioritised findings and fixes.
npx skillsauth add mohitagw15856/pm-claude-skills agent-design-reviewInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Most agents don't fail because the model is weak — they fail because the design lets them loop, call the
wrong tool, lose the thread across steps, or burn tokens with no stopping rule. This skill reviews an agent's
architecture against the decisions that actually determine reliability, and ranks the fixes — so "it works in
the demo but not in prod" becomes a specific list of changes. (Writing a new agent spec? Use
agent-spec.)
Given a sketch ("a research agent that searches, reads, and writes a report"), deliver the full review anyway — infer the likely control flow and tools, label the inference, and flag what to confirm. Never withhold the review for missing detail.
Ask for these only if they aren't already provided (else infer and label):
1. Summary — will this be reliable in production? The top 3 risks and the single change that helps most.
2. Findings by dimension — for each, what's sound and what's fragile:
| Dimension | Finding | Severity | Fix | |---|---|---|---| | Control flow | no max-steps / no progress check → loops | High | step budget + "am I making progress?" check + halt | | Tool use | overlapping tools confuse selection | Med | fewer, sharply-described tools; allowlist | | Context | full history re-sent each step → cost + drift | High | summarise/scope memory per step | | Failure handling | one tool error aborts the run | Med | retry/backoff + graceful degradation | | Safety | acts without confirmation on writes | High | human/confirm gate on side-effecting actions |
3. Reliability checklist — termination guarantee (it always stops), error recovery, idempotency of side-effecting actions, and determinism where it matters.
4. Cost & latency — where tokens/steps are spent and how to cut them (cheaper model for sub-steps, caching, fewer round-trips) without losing quality. Pair with llm-cost-latency-budget.
5. Safety — untrusted input/tool output handled as data not instructions, least-privilege tools, and
confirmation gates on high-impact actions. Pair with llm-guardrails-spec.
6. Prioritised fix plan — ordered by impact-to-effort.
LLM agent design practice — bounded control flow, least-privilege tool use, context management, error recovery, and safety gating.
business
Analyze why deals are won and lost and turn it into an action plan. Use when asked to run a win/loss analysis, review closed-won and closed-lost deals, understand why the team is losing to a competitor, or summarize sales feedback into patterns. Produces a structured win/loss report with themes, win/loss rates by segment and competitor, representative quotes, and prioritized actions for product, marketing, and sales.
development
Route a fuzzy request to the right skill in this library. Use when the user is unsure which skill fits, asks 'which skill should I use for X', describes a task without naming a skill, or when a request could plausibly match several skills. Produces a best-fit recommendation with the inputs to gather, a runner-up with the tie-breaker, and a workflow recipe when the job spans multiple skills.
testing
Triage a vulnerability or scanner finding — assess real severity, exploitability, and how urgently to fix. Use when asked to triage a CVE, prioritize scanner/pentest findings, assess a vuln's risk, or decide what to patch first. Produces a triage verdict: CVSS-informed severity adjusted for your context, exploitability, real risk, a fix/mitigation, and an SLA — so you fix what matters, not just what's red.
development
Stand up a Voice of Customer (VoC) program that turns feedback into action. Use when asked to build a VoC program, design a customer feedback loop, consolidate feedback sources, or set up a closed-loop feedback process. Produces a VoC program design — objectives, feedback sources and channels, a taxonomy, collection and analysis cadence, closed-loop routing, ownership, and success metrics.