skills/category-allocation-best-response/SKILL.md
Computes the best-response allocation of roster resources across categories in a Head-to-Head Categories matchup. Given our per-category capacity, the opponent's projected output, per-category win probabilities (from matchup-win-probability-sim), and a K-of-N winning threshold, classifies categories into pushed / contested / conceded buckets, emits per-category leverage weights for downstream lineup and streaming decisions, computes the resulting K-of-N win probability, and writes a plain-English rationale. Domain-neutral — portable to any fantasy sport with H2H Cats scoring (MLB 10-cat, NBA 9-cat, NHL 10-cat). Use when you need push/punt decisions, dominated-strategy elimination, leverage weights per cat, or best-response allocation; or when the user mentions "category allocation", "push or punt", "K of N cats", "dominated strategy elimination", "best response allocation", "Blotto fantasy", "leverage weights per cat", or "which cats to push".
npx skillsauth add lyndonkl/claude category-allocation-best-responseInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Scenario: Yahoo MLB 10-category H2H, Week 8. Threshold = 6 of 10. Upstream matchup-win-probability-sim returned per_cat_win_probability for the current week. Our resources: 5 open roster slots, $80 FAAB, 3 streamer starts.
Inputs:
| Cat | our_capacity | opp_projection | per_cat_win_prob | Inverse? | |-----|--------------|----------------|------------------|----------| | R | 42 | 38 | 0.64 | no | | HR | 12 | 14 | 0.35 | no | | RBI | 40 | 41 | 0.47 | no | | SB | 6 | 4 | 0.72 | no | | OBP | 0.335 | 0.328 | 0.64 | no | | K | 55 | 50 | 0.64 | no | | ERA | 3.85 | 4.10 | 0.65 | yes | | WHIP | 1.22 | 1.28 | 0.68 | yes | | QS | 4 | 3 | 0.69 | no | | SV | 2 | 5 | 0.08 | no |
Classification by win probability:
conceded — dominated strategy. Opponent has 5 projected SV to our 2; no marginal roster slot flips this. Leverage = 0.contested — borderline-losing but within reach. Leverage = 1.5.contested — coin-flip. Leverage = 1.5.pushed (lock-in needs attention, 0.25–0.85). Leverage = 1.2. (0.70 is the boundary — one of these would sit on the contested side at 1.5 if very close to the line.)pushed. Leverage = 1.2.No cat exceeds 0.85 on this week; no cats are locked.
Outputs:
pushed_cats: [SB, QS, WHIP, ERA, K, R, OBP] (7 cats at 0.64–0.72, ordered by win prob descending)conceded_cats: [SV]contested_cats: [HR, RBI]leverage_weights: {R: 1.2, HR: 1.5, RBI: 1.5, SB: 1.2, OBP: 1.2, K: 1.2, ERA: 1.2, WHIP: 1.2, QS: 1.2, SV: 0.0}k_of_n_win_probability: 0.605 (Poisson-binomial over the 10 per-cat probabilities with threshold 6)rationale: "SV is a dominated strategy (8% flip probability) — do not spend a reliever slot. The 7 pushed cats cover the 6-cat threshold on median, so lock them in and invest marginal resources into HR and RBI (contested, 1.5x leverage) where one extra power bat can flip the matchup."Interpretation: if we simply defend the 7 pushed cats we hit threshold in expectation. The contested cats (HR, RBI) are where the highest-leverage moves live — every marginal HR or RBI unit maps almost 1-for-1 to matchup win probability.
Copy this checklist and track progress:
Category Allocation Best-Response Progress:
- [ ] Step 1: Validate inputs and confirm upstream signals
- [ ] Step 2: Classify each cat by per_cat_win_probability
- [ ] Step 3: Assign leverage_weights
- [ ] Step 4: Check threshold satisfaction (pushed + contested >= K?)
- [ ] Step 5: Apply borderline-upgrade logic if threshold not met
- [ ] Step 6: Compute k_of_n_win_probability via Poisson-binomial
- [ ] Step 7: Write rationale and emit outputs
Step 1: Validate inputs and confirm upstream signals
per_cat_win_probability is the critical upstream dependency and must come from matchup-win-probability-sim (or an equivalent simulator). Do not invent per-cat probabilities from raw capacity vs projection — matchup-win-probability-sim accounts for variance, and variance is what determines flip probability for contested cats. See resources/methodology.md.
our_per_cat_capacity, opp_per_cat_projection, and per_cat_win_probabilityper_cat_win_probability values are in [0, 1]cat_win_threshold is in [1, N] where N = len(cats)inverse_cats is a subset of the cat listresources_available dict is present (even if empty — downstream skills may still consume leverage weights without a resource plan)random_seed from matchup-win-probability-sim is recorded for auditStep 2: Classify each cat by per_cat_win_probability
Apply the four-tier classification in Quick Reference. The thresholds are deliberate — see resources/methodology.md for the Blotto-derived rationale.
< 0.25 → conceded_cats (dominated strategy — any roster slot here is strictly worse than redeploying)0.25–0.70 → contested_cats (highest marginal value of one more unit)0.70–0.85 → pushed_cats (needs attention but should hold)> 0.85 → locked (do not waste marginal resources — returns are near-zero)per_cat_win_probability directly — the inverse handling has already been applied upstream by matchup-win-probability-sim. Do not re-invert.Step 3: Assign leverage_weights
Leverage weights propagate to mlb-lineup-optimizer, mlb-streaming-strategist, mlb-waiver-analyst, and any downstream consumer that maximizes Σ daily_quality × leverage[cat]. See resources/methodology.md.
leverage = 0.0 for conceded_cats (hard zero — not low, zero)leverage = 1.5 for contested_cats (high marginal value of one more unit)leverage = 1.2 for pushed_cats (above default but below contested)leverage = 1.0 for locked cats (default — marginal gains are redundant)Step 4: Threshold satisfaction check
Verify len(pushed_cats) + len(contested_cats) >= cat_win_threshold. If not, we are mathematically unable to reach the win threshold even if we go 100% on our defensible cats; the matchup is presumptively losing and we must either upgrade a borderline-conceded cat or pivot to variance-seeking play. See principle #6 in frameworks/game-theory-principles.md.
pushed_cats + contested_cats (these are the cats where we have nonzero flip probability)>= cat_win_threshold: threshold satisfied, proceed to Step 6< cat_win_threshold: apply Step 5 borderline-upgrade logicvariance-strategy-selector as a high-variance play candidateStep 5: Borderline-upgrade logic (conditional)
If Step 4 fails, look for conceded_cats with per_cat_win_probability in the [0.20, 0.25) "upgrade band". These are just below the concede line — a moderate resource investment (one waiver add, one FAAB bid, one streamer start) can lift them over 0.25 and into the contested tier.
conceded_cats by per_cat_win_probability descendingper_cat_win_probability >= 0.20contested and set leverage = 1.5Step 6: Compute k_of_n_win_probability via Poisson-binomial
Treat per-cat wins as independent Bernoullis (the same approximation used by matchup-win-probability-sim in poisson_binomial mode). Compute P(sum >= cat_win_threshold) via the standard PB recurrence. See resources/methodology.md.
P_0(0) = 1
P_i(k) = P_{i-1}(k) * (1 - p_i) + P_{i-1}(k-1) * p_i for i = 1..N, k = 0..N
k_of_n_win_probability = sum over k >= threshold of P_N(k)
per_cat_win_probability vector (Step 5 may have modified one entry)k_of_n_win_probabilitymatchup_win_probability returned by matchup-win-probability-sim by more than 0.03, investigate — the two should match within PB approximation errorStep 7: Write rationale and emit outputs
Rationale is a 2–4 sentence plain-English summary of the allocation logic. Name the conceded cats and why, name the contested cats and the highest-leverage resource move, call out any borderline upgrades, and state the computed k_of_n_win_probability. See resources/template.md for worked examples.
pushed_cats, conceded_cats, contested_cats, leverage_weights, k_of_n_win_probability, rationaleleverage_weights has one entry per cat (not one per push/concede/contest bucket)Pattern 1: MLB 10-cat (Yahoo 5x5)
[R, HR, RBI, SB, OBP, K, ERA, WHIP, QS, SV][ERA, WHIP]mlb-lineup-optimizer (principle #5), mlb-streaming-strategist (principle #1), mlb-waiver-analyst (principle #2)Pattern 2: NBA 9-cat
[PTS, REB, AST, STL, BLK, 3PM, FG%, FT%, TO][TO]pushed tier tends to move together, so a single roster move can lift multiple cats at once. This makes contested-cat leverage particularly high.Pattern 3: NHL 10-cat
[G, A, +/-, PIM, PPP, SOG, W, GAA, SV%, SO][GAA]Pattern 4: Variance-seeking underdog (cross-domain)
k_of_n_win_probability < 0.40: we are the underdog. Raise leverage on contested cats to 1.5, and flag the matchup for variance-strategy-selector. High-variance lineup construction (boom-bust players, one-start studs) can lift our win probability from ~35% to ~45% even without roster improvements. See principle #6 in frameworks/game-theory-principles.md.Do not invent per-cat probabilities. This skill is a downstream consumer of matchup-win-probability-sim. If you do not have per_cat_win_probability from a proper simulator, do not proceed — compute them first. Eyeballing probabilities from raw projections ignores variance, which is the whole point of the contested-cat classification.
Leverage weight 0.0 is a hard constraint, not a preference. When mlb-lineup-optimizer reads leverage[SV] = 0.0, it will refuse to start a pure-SV reliever even if the reliever has a high daily_quality. This is correct behavior (principle #1 in frameworks/game-theory-principles.md) — do not soften the zero to 0.1 or 0.2 to "keep options open". Zero means zero.
Contested-cat leverage stays 1.5 even if we are favored in the cat. The 1.5 multiplier captures the marginal value of one extra unit — which is high whenever the cat is close. A cat at per_cat_win_probability = 0.68 still has room for marginal gains to flip it to a near-certain win; leverage remains 1.5 (it is a pushed cat at 1.2, and moves to 1.5 if it drops into the contested band below 0.70).
Threshold and count must match the league format. cat_win_threshold = 6 for 10-cat MLB (strict majority). 5 for 9-cat NBA. 6 for 10-cat NHL. Passing the wrong threshold produces a silently wrong classification. Always confirm the league's tie-break rules — some leagues award ties for half-wins.
Dominated-strategy elimination is per-matchup, not per-season. A cat conceded this week vs a closer-heavy opponent may be a push cat next week vs a different opponent. Do not cache conceded_cats across weeks. Re-run the classification every matchup.
Inverse-cat handling is upstream's job. matchup-win-probability-sim already returns per_cat_win_probability with inverse handling applied (ERA 3.50 beats ERA 4.20 = win probability near 1.0). This skill consumes those probabilities directly. Do not re-invert, and do not treat inverse cats differently at the classification step.
Upgrade band is narrow (0.20–0.25). Only upgrade a conceded cat when its per_cat_win_probability is within the narrow [0.20, 0.25) band and we have resources available. Upgrading a 0.15-probability cat is throwing resources at a losing cause. Upgrading a 0.24-probability cat with a single waiver add may flip it to 0.30 — worth it.
Document resource cost when recommending upgrades. Abstract "upgrade HR" is unactionable. Say "upgrade HR with 1 roster slot + $15 FAAB on a power bat who adds ~3 HR/week", so mlb-waiver-analyst and mlb-faab-sizer have a concrete target.
When the matchup is presumptively losing, pivot to variance, not to pushing harder. If k_of_n_win_probability < 0.40 after classification and upgrades, the best move is to maximize variance (principle #6), not to redouble on pushed_cats. Flag the matchup in the rationale and call variance-strategy-selector downstream.
The skill is domain-neutral — resist MLB-specific assumptions. When called from NBA or NHL contexts, the thresholds (0.25 / 0.70 / 0.85), bands, leverage weights, and upgrade-band logic all apply unchanged. Only the cat list and win threshold change between sports. Keep the core logic sport-agnostic.
Four-tier classification (from per_cat_win_probability):
| Range | Bucket | Leverage | Rationale |
|-------|--------|----------|-----------|
| [0.00, 0.25) | conceded | 0.0 | Dominated strategy. Marginal unit has near-zero flip impact. |
| [0.25, 0.70) | contested | 1.5 | Highest marginal value — one more unit often flips outcome. |
| [0.70, 0.85] | pushed | 1.2 | Should hold but not guaranteed; defend with attention. |
| (0.85, 1.00] | locked | 1.0 | Default weight; marginal gains are near-redundant. |
Upgrade band (for borderline concedes):
| Range | Action |
|-------|--------|
| [0.20, 0.25) | Candidate for upgrade if Step 4 threshold check fails. |
| [0.00, 0.20) | Do not upgrade. Resources are better spent on contested cats. |
Poisson-binomial recurrence (for k_of_n_win_probability):
Given p = [p_1, p_2, ..., p_N] and threshold K:
P_0(0) = 1
P_0(k) = 0 for k >= 1
For i = 1..N:
P_i(0) = P_{i-1}(0) * (1 - p_i)
For k = 1..i:
P_i(k) = P_{i-1}(k) * (1 - p_i) + P_{i-1}(k-1) * p_i
k_of_n_win_probability = sum_{k=K}^{N} P_N(k)
Inputs required:
our_per_cat_capacity: dict[cat, number] — our projected per-cat output (remaining-week or full-week)opp_per_cat_projection: dict[cat, number] — opponent's projected per-cat outputper_cat_win_probability: dict[cat, float in [0,1]] — FROM matchup-win-probability-simcat_win_threshold: int — 6 for MLB 10-cat, 5 for NBA 9-cat, 6 for NHL 10-catresources_available: dict[str, number] — e.g., {"roster_slots": 5, "faab": 80, "streamer_starts": 3}inverse_cats: list[str] — cats where lower is better (ERA, WHIP, TO, GAA, etc.)Outputs produced:
pushed_cats: list[cat] — ordered by per_cat_win_probability descendingconceded_cats: list[cat] — dominated; leverage 0.0contested_cats: list[cat] — highest marginal leverageleverage_weights: dict[cat, float] — values in {0.0, 1.0, 1.2, 1.5}k_of_n_win_probability: float in [0,1] — overall matchup win prob under this allocationrationale: string — 2–4 sentences of plain-English allocation logicUpstream dependencies:
matchup-win-probability-sim (REQUIRED) — supplies per_cat_win_probability and overall matchup_win_probability for cross-validation*-category-state-analyzer (optional, MLB/NBA/NHL) — supplies the per-cat capacity and opponent projectionDownstream consumers:
mlb-lineup-optimizer / NBA / NHL equivalents — consume leverage_weights (principle #5)mlb-streaming-strategist — reads conceded_cats as hard constraint (principle #1)mlb-waiver-analyst / mlb-faab-sizer — consume contested_cats + resources_available to target high-leverage addsvariance-strategy-selector — consumes k_of_n_win_probability to decide favorite-vs-underdog play style (principle #6)Key resources:
testing
Cluster a conference's event records into a small set of coarse themes with finer sub-clusters, an explicit outlier bucket, and soft (multi-membership) affinities — using the hybrid embed-then-label pipeline (embed abstracts, reduce, density-cluster, then LLM-label the clusters) when embedding libraries are available, and an LLM-reasoned hierarchical fallback when they are not. Embeddings do the grouping; the LLM only names the groups. Conference-agnostic. Use when turning structured event records into a navigable theme map for preference elicitation and scheduling, when you need 6-8 reasonable themes rather than 20 muddy ones, or when overlapping talks must belong to more than one theme. Trigger keywords - theme clustering, cluster talks, embed then label, soft membership, outlier talks, conference themes, topic map.
development
Build a personal conference schedule as a constraint-optimization problem — hard constraints (no time overlap, room-to-room travel time, capacity/registration, the attendee's own must-attends and blackouts) plus a user-owned weighted objective trading interest against breadth, pacing (maximize contiguous free time), and serendipity. Surfaces unbreakable conflicts (two high-value overlapping talks the model cannot rank) as decisions for the human rather than silently picking, and reports what each choice traded away. Conference-agnostic. Use to turn a preference profile plus a theme map into a day-by-day plan, to resolve overlapping sessions, or to balance a packed vs paced schedule. Trigger keywords - schedule optimization, conference schedule, constraint optimization, overlapping talks, contiguous free time, conflict surfacing, packed vs paced.
development
Parse a heterogeneous conference program (markdown, HTML, PDF-derived text, or JSON) into normalized event records with per-field confidence scores and independent classification axes (topic, depth, format, prerequisites, recorded, capacity). Detects the program's format before extracting, treats every inferred field as uncertain (present vs inferred vs missing), and flags thin or missing abstracts so downstream enrichment can target them. Conference-agnostic. Use when ingesting a conference or event schedule into a structured store, normalizing a talk/session list, or extracting per-session metadata with calibrated confidence. Trigger keywords - program ingestion, parse schedule, session extraction, event records, conference program, talk metadata, per-field confidence.
development
Build a personalized preference profile from a small number of well-chosen, cluster-grounded questions instead of a long survey. Represents the person's interests as an uncertainty region over the theme map, picks the single highest-information-gain choice-based question (contrasting real talks from different clusters), balances exploiting known interests against exploring uncertain ones, deliberately injects outlier probes to fight selection bias, and stops as soon as the schedule would be stable. Also elicits the user-owned objective weights and hard constraints. Interactive — runs where it can actually ask the person. Conference-agnostic. Use to turn a theme map into a preference profile, to decide what to ask a conference attendee, or to elicit scheduling priorities. Trigger keywords - preference elicitation, ask few questions, information gain, choice-based questions, selection bias probe, objective weights, attendee preferences.