skills/prompt-regression-suite/SKILL.md
Design a regression test suite that catches an LLM feature getting worse when the prompt, model, or context changes. Use when asked to stop prompt changes breaking production, set up golden tests or CI gates for an LLM feature, or test a model/prompt upgrade before shipping it. Produces a golden case set, per-case pass criteria, CI gate thresholds, and a triage protocol for failures. For designing first-time evaluation of a new feature use ai-eval-plan instead.
npx skillsauth add mohitagw15856/pm-claude-skills prompt-regression-suiteInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Every prompt tweak, model upgrade, and context change is a deploy. This skill designs the suite that runs on each one and answers a single question: did anything that used to work stop working?
Ask for (if not already provided):
Compose the set from four deliberate classes — not a random sample:
| Class | Purpose | Share | |---|---|---| | Core paths | The 5-10 inputs that represent most real traffic | ~40% | | Past failures | Every input that caused a bug, complaint, or incident — permanently | ~25% | | Edge & adversarial | Empty/huge inputs, wrong language, injection attempts, off-topic | ~25% | | Canaries | Cases pinned to behaviours you never want to change (refusals, format, tone) | ~10% |
Keep it small enough to run on every change (30-80 cases beats 500 nobody runs). Version it in git next to the prompt.
Choose the cheapest check that catches the regression:
Every case records: input, pass criteria, scoring method, and the baseline output at the time it was added.
llm-cost-latency-budget).When a case fails, classify before "fixing":
Trigger: runs on [prompt edit / model bump / retrieval change] via [CI job].
Golden set ([n] cases):
| # | Class | Input (summary) | Pass criteria | Method | |---|---|---|---|---|
Gates: merge blocks when [conditions]. Warnings on [conditions].
Triage: [the three-way protocol, with who owns updates to the golden set]
Maintenance: every production incident adds a case within [period]; the set is reviewed for dead cases each [quarter].
business
Analyze why deals are won and lost and turn it into an action plan. Use when asked to run a win/loss analysis, review closed-won and closed-lost deals, understand why the team is losing to a competitor, or summarize sales feedback into patterns. Produces a structured win/loss report with themes, win/loss rates by segment and competitor, representative quotes, and prioritized actions for product, marketing, and sales.
development
Route a fuzzy request to the right skill in this library. Use when the user is unsure which skill fits, asks 'which skill should I use for X', describes a task without naming a skill, or when a request could plausibly match several skills. Produces a best-fit recommendation with the inputs to gather, a runner-up with the tie-breaker, and a workflow recipe when the job spans multiple skills.
testing
Triage a vulnerability or scanner finding — assess real severity, exploitability, and how urgently to fix. Use when asked to triage a CVE, prioritize scanner/pentest findings, assess a vuln's risk, or decide what to patch first. Produces a triage verdict: CVSS-informed severity adjusted for your context, exploitability, real risk, a fix/mitigation, and an SLA — so you fix what matters, not just what's red.
development
Stand up a Voice of Customer (VoC) program that turns feedback into action. Use when asked to build a VoC program, design a customer feedback loop, consolidate feedback sources, or set up a closed-loop feedback process. Produces a VoC program design — objectives, feedback sources and channels, a taxonomy, collection and analysis cadence, closed-loop routing, ownership, and success metrics.