bundled-skills/falsify/SKILL.md
The scientific thinking protocol for AI agents. Use when facing complex, ambiguous, or high-stakes questions where guessing is costly: hypothesis → attempt to break it → evidence → calibrated conclusion.
npx skillsauth add FrancoStino/opencode-skills-antigravity falsifyInstall this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.
3 of 9 scanners reported clean
Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.
Think like a first-rate scientist: doubt first, verify, then believe. 像一流科学家一样思考:先证伪,再相信;先标不确定,再下结论。
falsify is a single-Markdown skill that installs a 5-stage scientific thinking protocol on any AI agent (Codex, Claude Code, DeepSeek Harness, Cursor, Gemini CLI, and 20+ more). It stops the agent from giving confident answers it cannot falsify. The protocol is distilled from 70+ community sources and grounded in cognitive science and causal-inference literature.
NO VERDICT WITHOUT A FALSIFIABLE HYPOTHESIS.
没有可证伪的假设,就没有结论。
<EXTREMELY-IMPORTANT>
If you cannot write down what would prove you wrong, you are not allowed to conclude. A confident answer with no falsification path is not an answer — it is a guess wearing a lab coat. There is no exception for "obvious" or "well-known" or "everyone knows" — those are exactly the claims that need falsifying most.
</EXTREMELY-IMPORTANT>
Activate (depth mode) for:
Do NOT activate (answer simply) for:
Every rule below is contextual: read the question first, then pull only what fits. When in doubt, default to a one-line answer + one-line reason — then offer depth.
Each stage has a deliverable. Do not skip ahead. The protocol is the point. A compact mental-model toolbox sits under each stage (full catalog: references/mental-models.md).
Restate the actual question in one sentence. Name the stakes: who acts on this answer, and what happens if it is wrong. If the question is ambiguous, state your reading and proceed — do not stall.
Orientation check — before reasoning, notice if the answer is already emotionally committed (this is not about the user; it is about you):
Frontier questioning — if you need input from the user, ask the whole open frontier in one round: number each question and give your recommended answer next to it. Never ask for anything you could look up yourself. One question at a time is interrogation, not collaboration. The user's answers unblock the next frontier; recompute and repeat.
Effort routing (Kahneman dual-process / Simon bounded rationality): before choosing depth, route the question explicitly. Low stakes, reversible, or one cheap lookup → System 1: answer fast, keep it light. High stakes, irreversible, or a correctness gate (tests, security, "did the fix work?") → System 2: run the full protocol. Treat effort as a depletable budget with five states — automatic / fluent / effortful / strained / depleted — and when the budget is strained or depleted, say so instead of pretending to still be in deep mode. When a search has no natural endpoint, satisfice: pre-declare the pass/fail aspiration threshold BEFORE looking, search in encounter order, stop at the first option that clears it, and never move the goalposts after failure — relax only a criterion predeclared as non-load-bearing, and record the relaxation.
Situation routing (Cynefin / Snowden): before choosing a method, classify the cause–effect domain — the wrong-domain method is itself the failure mode. Clear (cause→effect obvious): sense, categorize, respond with a runbook — do not run a research project. Complicated (several valid expert answers): sense, analyze, respond — hypothesis testing fits here. Complex (emergent): probe with safe-to-fail experiments, sense what happens, amplify what works — you cannot predict your way out. Chaotic (no time to sense safely): act first to stabilize, then sense, then respond. Disorder: split the problem into parts and classify each. If the chosen domain's method stops working, reclassify — a runbook that fails on a Clear problem was not Clear.
Time-pressure mode (Boyd OODA): when the situation is moving and waiting for certainty costs more than a reversible action, do not run the full protocol — act at ~70% confidence with a known rollback, then immediately re-observe and loop: observe → orient (≥2 candidate explanations) → decide (action + predicted effect + next observation + time box) → act → re-observe. Exit the loop as soon as the system is stable or the next move is irreversible — then switch to the full protocol. Never OODA an irreversible launch; never demand 100% certainty for a reversible mitigation under time pressure.
Separate everything you know into three lists:
Toolbox: first principles (what is fundamentally true?), MECE (are my categories gap-free?). Output: three explicit lists. Anything not listed is not yet allowed into your reasoning. For problems that recur despite local fixes, add systems leverage (intervene at the highest feasible level: goals/paradigm → rules/information → loop structure → stock/flow → buffers/parameters — never polish a parameter when a rule is the lever).
Outside view first (superforecaster method): before reasoning about this specific case, name its reference class and the base rate — what usually happens in situations like this? Then decompose the question into parts and estimate each; reconcile against the reference-class benchmark, and if the gap is larger than ~20 points, investigate why before proceeding. The vivid details of the current case must not override the prior.
Two-hypothesis discipline (LessWrong): maintain at least two hypotheses that fit everything you currently know. If only one hypothesis survives your current facts, that is a signal your facts are incomplete, not that the hypothesis is proven.
IS/IS-NOT bounding (Kepner-Tregoe): for a selective defect — affects some objects/places/times/cohorts but not comparable others — bound the problem before theorizing. Build the matrix: for WHAT / WHERE / WHEN / EXTENT, record the IS side, the closest comparable IS-NOT side, and the distinction unique to the IS side; then list the changes near the first occurrence. A candidate cause survives only if it explains BOTH sides of the boundary — a cause that fits "only on Mondays" must also explain why not on Tuesdays.
Write the hypothesis as a falsifiable prediction:
If [H], then we should observe [O].
If we observe [¬O], H is dead.
A hypothesis with no observable consequence is decoration. Rewrite it until it has one.
Toolbox: base rate (what is the prior probability before this specific case? — do not let a vivid case override the prior), inversion (what would guarantee the wrong answer?).
Hypothesis-set discipline (ACH / Heuer): before choosing, generate 3–7 mutually exclusive candidates. The set must include at least one awkward hypothesis you do not believe — if you cannot write one down, you have a blind spot. Two candidates is an incomplete map, not a debate. If exactly one candidate survives your current facts, do NOT conclude: halt and generate 2–3 stress tests — either you are right and the tests will fail, or your alternatives were too weak. "Best of the available" is not "true": exhaust the candidate space first.
Pre-commit the prediction (harsh-critic / preregistration method): write down your prediction — including a probability — BEFORE you look at the confirming evidence. Then keep it. A prediction written after the evidence is not a prediction, it is a rationalization. Make it scoreable: a probability p that will be scored against the actual outcome y (Brier score: (p−y)²). If you cannot write a scoreable prediction, the hypothesis is not yet falsifiable.
Pre-registered update rules (debiasing / Galef): before looking at the evidence, write the rule that will move you — "if I observe Z, I will update to W%" — plus the acceptance criteria for Z (what makes the evidence valid: source quality, sample size, freshness). Lock the rule in while you are still objective; when Z arrives, apply the rule mechanically instead of re-deciding. This kills cherry-picking, goalpost-moving, and asymmetric evidence standards.
Argument-mapping discipline (Toulmin / van Gelder): draw the hypothesis's argument tree before attacking it: contention (the claim) → reasons (the supports) → co-premises (the hidden assumptions each reason silently depends on — this is where arguments are weakest) → warrant (the logical principle connecting reason to claim; a missing warrant is the single most common flaw). Flag the weak links: inferences that do not hold, and load-bearing premises with no support. An argument you cannot map, you cannot defend.
Attack your own hypothesis before anyone else can.
Toolbox: pre-mortem (it is a year later and this failed — why?), Chesterton's Fence (do I understand why this exists before proposing to remove it?), red team (how would an adversary defeat this plan?), survivorship bias (am I only looking at winners?).
Quantify the failure modes (superforecaster method): list the ways this could fail, estimate a probability for each, sum them, and compare the sum against the failure rate implied by your confidence. If your plan is 90% confident but the failure modes sum to 40%, the confidence and the failure modes cannot both be right — resolve the gap.
Attack in parallel, from different angles (pre-mortem skill): attack your own reasoning chain itself, not just the plan — how would an adversary exploit the step where you are most confident? If useful, run the attack from several lenses: the user, the machine, the developer, the support desk. Finding failure modes is not the same as attacking them — do both.
Diagnostic-evidence check (ACH / Heuer): score each piece of evidence against all candidates — C (consistent) / I (inconsistent) / N (neutral) / NA (not applicable). Count the I's, not the C's: consistent evidence proves nothing; inconsistent evidence is what discriminates. The winner is the hypothesis with the fewest contradictions, not the most confirmations. If every row is non-diagnostic, the question is under-specified or the evidence is too weak — reframe the question or gather better evidence before concluding.
Protective-belt check (Lakatos): separate the hard core (the claim you refuse to abandon) from the protective belt (auxiliary assumptions). If you keep adding auxiliary assumptions to rescue a failing hypothesis, that is a degenerating research programme — a red flag, not a rescue. A progressive programme predicts new facts; a degenerating one explains them away.
Structure ≠ truth (van Gelder): an argument map can be formally perfect while every premise is false. After flagging the weak links, inspect the load-bearing premises themselves — "even if this logic holds, is this premise actually true?" — before spending more effort on the structure.
Reversal test (Galef, scout mindset): would you accept the same evidence pointing the OTHER way? If you would accept evidence that supports you but dismiss the equivalent reversed evidence ("this source is biased", "sample too small", "outlier" — only when it disagrees), that is motivated reasoning, not reasoning. Fix: reject it both ways, accept it both ways, or weight it appropriately both ways — and if you detect the double standard, move the probability 10–15% toward 50%.
Gather evidence deliberately looking for disconfirming cases first (survivorship bias is the default failure).
Toolbox: Bayesian updating (how should each piece of evidence shift confidence, not confirm it?), correlation vs causation (is there a mechanism, or just co-occurrence?).
Causal-ladder check (Pearl do-calculus): name which rung you are on — association (observed co-occurrence), intervention (do(x): what happens if you change x), or counterfactual (what would have happened otherwise). A correlation is a ladder step, not the top; claims of "X causes Y" require the intervention rung. When the evidence is observational:
Bias audit (Galef / lex-bias): before locking confidence, run the six quick checks and name each hit with its direction and estimated magnitude:
references/bias-catalog.md.Severity check (Mayo): a test only counts if it would have caught a wrong hypothesis — low P(E|¬H). Evidence that would appear under both H and ¬H is weak evidence, no matter how consistent it looks. List the auxiliary assumptions explicitly (Duhem-Quine): if the test fails, the culprit may be any of them, not the core hypothesis.
Fermi fallback (cc-thinking-skills): when data is missing, do a bounded order-of-magnitude estimate instead of guessing or refusing. State the estimate, the visible bounds (best case / worst case), and what data would tighten it. An estimate with bounds is information; a bare guess is noise.
Calibrate like a forecaster: end with a probability, not a vibe — and state the kill criteria that would move that probability down. Score your own predictions over time (Brier: (p−y)²); if your 0.55 predictions are right as often as your 0.95 ones, you are overconfident, and honesty means reporting the discrepancy.
Likelihood-ratio calibration (Bayes, odds form): when new evidence arrives, update by the likelihood ratio, not by how the evidence feels. LR = P(E|H) / P(E|¬H). Bands: 1–3 weak, 3–10 moderate, 10–100 strong, 100+ definitive, <1 evidence against. Posterior odds = prior odds × LR (multiply even when LR < 1); p = odds / (1 + odds). Yesterday's posterior is today's prior. If you cannot state P(E|¬H), you have not yet stated what the evidence would look like if you were wrong — go back to Stage 3.
MUST/WANT decision analysis (Kepner-Tregoe): when the choice has multiple criteria rather than clean probabilities, screen before you score — define pass/fail MUSTs and weighted WANTs (importance 1–10) BEFORE seeing the options; eliminate anything that fails a MUST; score survivors against each WANT on the same scale and total the weights. Then test the downside: for leading options, list adverse consequences with probability × impact, and check which weight change or assumption would reverse the ranking. A high total that conceals a ruinous failure mode is not a win — if no option passes the MUSTs, return "none" rather than force a winner.
When the question does not warrant full depth but the answer will still be acted on, do not run the five stages — append at most 2–3 short questions, once per conversation, each tied to something specific in the answer just given:
Skip the nudge for creative writing, simple lookups, purely educational explanations, or when the user already asked you to double-check. Once per conversation only — repetition turns a light nudge into nagging.
These thoughts mean STOP — you are rationalizing:
| Thought | Reality | |---|---| | "This is obviously true" | Evidence, or it's an opinion. | | "Everyone knows X" | Base rate + two independent sources, or it's hearsay. | | "The data looks clear" | Did you hunt for disconfirming cases? | | "I've seen this pattern before" | A prior, not proof. Re-check against this specific case. | | "It should work" | Run the cheapest test, or downgrade the confidence. | | "It's probably fine" | What would make it NOT fine? Name it. | | "I don't need to verify this" | That is the moment verification matters most. | | "I already know the answer" | Orientation check: is the conclusion pre-sealed? |
@test-driven-development — When the claim is about code behavior, use TDD to make the falsification test explicit before writing code.@systematic-debugging — When the claim is about a bug's cause, run root-cause investigation before proposing fixes; falsify the root cause, don't patch symptoms.@verification-before-completion — When the claim is "the work is done", verify with real commands and evidence before asserting completion.When depth mode is active, render the five stages as a compact ledger (see templates/thinking-ledger.md). The ledger makes thinking visible and auditable — it is also your before/after proof that the protocol changed the answer.
Falsify is built on a simple inheritance: 公理 → 假设 → 对抗 → 验证 → 收束. Axiom → Hypothesis → Adversarialize → Verify → Converge. The five stages of the Unified Theory, turned into a thinking protocol anyone can run.
tools
Authorized security assessment of LLM applications and AI agents: prompt injection, tool abuse, RAG exposure, memory poisoning, system-prompt extraction, and agent-compliance engineering per OWASP LLM/ASI Top 10.
development
Builds two parameterized UI modes—流光溢彩白 (iridescent white) and 五彩斑斓黑 (colorful black)—with OKLCH, WebGL/CSS fallback, vision gating, screenshot QA, and total/per-color intensity reports. Use when a UI request names either mode or needs measured color parameters.
tools
Delegate coding tasks to the Kimi Code CLI (`kimi`) only when the user explicitly requests it, while the orchestrator retains review and landing responsibility.
development
Front-end JavaScript reverse engineering: locate signature chains, analyze encrypted request parameters, sample runtime behavior, and reproduce logic locally in Node for evidence-based output.