library/specializations/product-management/skills/ab-test-design/SKILL.md
Statistical experiment design and analysis capabilities for product experimentation
npx skillsauth add a5c-ai/babysitter A/B Test DesignInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Specialized skill for statistical experiment design and analysis capabilities. Enables product teams to design rigorous experiments, calculate sample sizes, and interpret results with statistical confidence.
This skill integrates with the following processes:
product-market-fit.js - Validation experiments for PMF hypothesesconversion-funnel-analysis.js - Funnel optimization experimentsbeta-program.js - A/B testing during beta phases{
"type": "object",
"properties": {
"experimentType": {
"type": "string",
"enum": ["ab", "multivariate", "sequential", "bandit"],
"description": "Type of experiment to design"
},
"hypothesis": {
"type": "string",
"description": "Hypothesis to test"
},
"primaryMetric": {
"type": "object",
"properties": {
"name": { "type": "string" },
"baseline": { "type": "number" },
"mde": { "type": "number", "description": "Minimum detectable effect" }
}
},
"guardrailMetrics": {
"type": "array",
"items": { "type": "string" },
"description": "Metrics that should not regress"
},
"trafficAllocation": {
"type": "number",
"description": "Percentage of traffic for experiment"
},
"confidenceLevel": {
"type": "number",
"default": 0.95,
"description": "Statistical confidence level"
}
},
"required": ["experimentType", "hypothesis", "primaryMetric"]
}
{
"type": "object",
"properties": {
"experimentPlan": {
"type": "object",
"properties": {
"name": { "type": "string" },
"hypothesis": { "type": "string" },
"variants": { "type": "array", "items": { "type": "object" } },
"sampleSize": { "type": "number" },
"duration": { "type": "string" },
"metrics": { "type": "object" }
}
},
"powerAnalysis": {
"type": "object",
"properties": {
"requiredSampleSize": { "type": "number" },
"estimatedDuration": { "type": "string" },
"power": { "type": "number" }
}
},
"implementation": {
"type": "object",
"properties": {
"trackingEvents": { "type": "array", "items": { "type": "string" } },
"segmentation": { "type": "array", "items": { "type": "string" } },
"rolloutPlan": { "type": "string" }
}
},
"analysisFramework": {
"type": "object",
"properties": {
"primaryAnalysis": { "type": "string" },
"secondaryAnalyses": { "type": "array", "items": { "type": "string" } },
"decisionCriteria": { "type": "object" }
}
}
}
}
const experimentDesign = await executeSkill('ab-test-design', {
experimentType: 'ab',
hypothesis: 'Adding social proof to pricing page increases conversion by 10%',
primaryMetric: {
name: 'pricing_page_conversion',
baseline: 0.05,
mde: 0.10
},
guardrailMetrics: ['revenue_per_visitor', 'bounce_rate'],
trafficAllocation: 50,
confidenceLevel: 0.95
});
development
Model documentation skill for generating model cards following Google's model card framework.
development
MLflow integration skill for experiment tracking, model registry, and artifact management. Enables LLMs to log experiments, compare runs, manage model lifecycle, and retrieve artifacts through the MLflow API.
data-ai
LIME-based local explanation skill for individual predictions across tabular, text, and image data.
devops
Kubeflow Pipelines skill for ML workflow orchestration, component management, and Kubernetes-native ML.