skills/ab-test-setup/SKILL.md
Structured guide for setting up A/B tests with mandatory gates for hypothesis, metrics, and execution readiness.
npx skillsauth add Regtransfers/agency-agents-mcp ab-test-setupInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
@ A/B Test Setup
@ 1️⃣ Purpose & Scope
Ensure every A/B test is valid, rigorous, and safe before a single line of code is written.
@ 2️⃣ Pre-Requisites
must have:
@ Hypothesis Quality Checklist
A valid hypothesis includes:
@ 3️⃣ Hypothesis Lock (Hard Gate)
Before designing variants or metrics, must:
Ask explicitly:
“Is this the final hypothesis we are committing to for this test?”
never proceed until confirmed.
@ 4️⃣ Assumptions & Validity Check (Mandatory)
Explicitly list assumptions about:
If assumptions are weak or violated:
@ 5️⃣ Test Type Selection
Choose the simplest valid test:
Default to A/B unless there is a clear reason otherwise.
@ 6️⃣ Metrics Definition
@ Primary Metric (Mandatory)
@ Secondary Metrics
@ Guardrail Metrics
@ 7️⃣ Sample Size & Duration
Define upfront:
Estimate:
never proceed without a realistic sample size estimate.
@ 8️⃣ Execution Readiness Gate (Hard Stop)
You may proceed to implementation only if all are true:
If any item is missing, stop and resolve it.
@ Running the Test
@ During the Test
DO:
never:
@ Analyzing Results
@ Analysis Discipline
When interpreting results:
@ Interpretation Outcomes
Result; Action
Significant positive; Consider rollout Significant negative; Reject variant, document learning Inconclusive; Consider more traffic or bolder change Guardrail failure; never ship, even if primary wins
@ Documentation & Learning
@ Test Record (Mandatory)
Document:
Store records in a shared, searchable location to avoid repeated failures.
@ Refusal Conditions (Safety)
Refuse to proceed if:
Explain why and recommend next steps.
@ Key Principles (Non-Negotiable)
@ Final Reminder
A/B testing is not about proving ideas right. It is about learning the truth with confidence.
If you feel tempted to rush, simplify, or “just try it” — that is the signal to slow down and re-check the design.
@ When to Use This skill is applicable to execute the workflow or actions described in the overview.
@ Limitations
tools
Build AI agents that interact with computers like humans do - viewing screens, moving cursors, clicking buttons, and typing text. Covers Anthropic's Computer Use, OpenAI's Operator/CUA, and open-source alternatives.
testing
Generate structured PR descriptions from diffs, add review checklists, risk assessments, and test coverage summaries. Use when the user says "write a PR description", "improve this PR", "summarize my changes", "PR review", "pull request", or asks to document a diff for reviewers.
tools
Use when working with comprehensive review full review
development
You are an expert in creating competitor comparison and alternative pages. Your goal is to build pages that rank for competitive search terms, provide genuine value to evaluators, and position your product effectively.