Adoption

Agent Skills are supported by leading AI development tools.

VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory

shaun-z/result-to-claim

Name: result-to-claim
Author: shaun-z

skills/skills-codex/result-to-claim/SKILL.md

npx skillsauth add shaun-z/auto-claude-code-research-in-sleep result-to-claim

Clean

TrivyContainer and dependency vulnerability scanner

Clean

SemgrepStatic code analysis for vulnerabilities

Clean

mcp-scan (Snyk)Model Context Protocol security validation

Skipped

Snyk (dep)Open source security scanning

Skipped

Socket.devSupply chain security analysis

Skipped

VirusTotalMulti-engine malware detection

Skipped

CrowdStrikeAdvanced threat intelligence

Skipped

OSV-ScannerOpen Source Vulnerability database check

Skipped

OWASP Dep-Check

Result-to-Claim Gate

Experiments produce numbers; this gate decides what those numbers mean. Collect results from available sources, get an objective judgment, then route based on the verdict.

Context: $ARGUMENTS

Constants

REVIEWER_MODEL = gpt-5.4 - Used via a secondary Codex agent for objective claim assessment.

When to Use

After a set of experiments completes (main results, not just sanity checks)
Before committing to claims in a paper or review response
When results are ambiguous and you need an objective second opinion

Workflow

Step 1: Collect Results

Gather experiment data from whatever sources are available in the project:

W&B (preferred): wandb.Api().run("<entity>/<project>/<run_id>").history() - metrics, training curves, comparisons
EXPERIMENT_LOG.md - full results table with baselines and verdicts
EXPERIMENT_TRACKER.md - check which experiments are done vs still running
Log files - ssh server "tail -100 /path/to/training.log" if no other source
docs/research_contract.md or project notes - intended claims and experiment design

Assemble the key information:

What experiments were run (method, dataset, config)
Main metrics and baseline comparisons (deltas)
The intended claim these experiments were designed to test
Any known confounds or caveats

Step 2: Secondary Codex Judgment

Send the collected results to a secondary Codex agent for objective evaluation:

spawn_agent:
  model: REVIEWER_MODEL
  reasoning_effort: xhigh
  message: |
    RESULT-TO-CLAIM EVALUATION

    I need you to judge whether experimental results support the intended claim.

    Intended claim: [the claim these experiments test]

    Experiments run:
    [list experiments with method, dataset, metrics]

    Results:
    [paste key numbers, comparison deltas, significance]

    Baselines:
    [baseline numbers and sources - reproduced or from paper]

    Known caveats:
    [any confounding factors, limited datasets, missing comparisons]

    Please evaluate:
    1. claim_supported: yes | partial | no
    2. what_results_support: what the data actually shows
    3. what_results_dont_support: where the data falls short of the claim
    4. missing_evidence: specific evidence gaps
    5. suggested_claim_revision: if the claim should be strengthened, weakened, or reframed
    6. next_experiments_needed: specific experiments to fill gaps (if any)
    7. confidence: high | medium | low

    Be honest. Do not inflate claims beyond what the data supports.
    A single positive result on one dataset does not support a general claim.

If delegation is unavailable, run the same evaluation locally and mark the verdict [pending external review] instead of blocking the pipeline.

Step 3: Parse and Normalize

Extract structured fields from the response:

- claim_supported: yes | partial | no
- what_results_support: "..."
- what_results_dont_support: "..."
- missing_evidence: "..."
- suggested_claim_revision: "..."
- next_experiments_needed: "..."
- confidence: high | medium | low

Step 4: Route Based on Verdict

`no` - Claim not supported

Record a postmortem in findings.md:
- What was tested, what failed, and hypotheses for why
- Constraints for future attempts (what not to try again)
Update the project pipeline status in project notes
Decide whether to pivot to the next idea from IDEA_CANDIDATES.md or try an alternative approach

`partial` - Claim partially supported

Update the working claim to reflect what is supported
Record the gap in findings.md
Design and run supplementary experiments to fill evidence gaps
Re-run /result-to-claim after supplementary experiments complete
If the same claim gets multiple partial verdicts, record the analysis in findings.md and consider narrowing the claim scope or switching ideas

`yes` - Claim supported

Record the confirmed claim in project notes
If ablation studies are incomplete, trigger /ablation-planner
If all evidence is in, move to paper writing

Rules

The secondary Codex agent is the judge, not the local executor. The local executor collects evidence and routes; the reviewer agent evaluates. This prevents post-hoc rationalization.
Do not inflate claims beyond what the data supports. If the verdict says partial, do not round up to yes.
A single positive result on one dataset does not support a general claim. Be honest about scope.
If confidence is low, treat the judgment as inconclusive and add experiments rather than committing to a claim.
If reviewer delegation is unavailable, make the best local judgment you can and mark it [pending external review].
Always record the verdict and reasoning in findings.md, regardless of outcome.

shaun-z/result-to-claim

skills/skills-codex/result-to-claim/SKILL.md

Use when experiments complete to judge what claims the results support, what they do not, and what evidence is still missing. A secondary Codex agent evaluates results against intended claims and routes to the next action (pivot, supplement, or confirm). Use after experiments finish - before writing the paper or running ablations.

development

Updated Apr 17, 2026

$ install --global

skillsauth

npx skillsauth add shaun-z/auto-claude-code-research-in-sleep result-to-claim

Install this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.

Security Scan Results

3 of 9 scanners reported clean

Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.

Scanners Passed

Scanners in report

Clean

TrivyContainer and dependency vulnerability scanner

95%

Clean

SemgrepStatic code analysis for vulnerabilities

95%

Clean

mcp-scan (Snyk)Model Context Protocol security validation

95%

Skipped

Snyk (dep)Open source security scanning

50%

Skipped

Socket.devSupply chain security analysis

50%

Skipped

VirusTotalMulti-engine malware detection

50%

Skipped

CrowdStrikeAdvanced threat intelligence

50%

Skipped

OSV-ScannerOpen Source Vulnerability database check

50%

Skipped

OWASP Dep-Check

50%

Last scanned: Mar 28, 2026, 7:25 AM36.2s1 file scanned

SKILL.md

name:: result-to-claim
description:: Use when experiments complete to judge what claims the results support, what they do not, and what evidence is still missing. A secondary Codex agent evaluates results against intended claims and routes to the next action (pivot, supplement, or confirm). Use after experiments finish - before writing the paper or running ablations.
allowed-tools:: Bash(*), Read, Grep, Glob, Write, Edit, Agent

Result-to-Claim Gate

Experiments produce numbers; this gate decides what those numbers mean. Collect results from available sources, get an objective judgment, then route based on the verdict.

Context: $ARGUMENTS

Constants

REVIEWER_MODEL = gpt-5.4 - Used via a secondary Codex agent for objective claim assessment.

When to Use

After a set of experiments completes (main results, not just sanity checks)
Before committing to claims in a paper or review response
When results are ambiguous and you need an objective second opinion

Workflow

Step 1: Collect Results

Gather experiment data from whatever sources are available in the project:

W&B (preferred): wandb.Api().run("<entity>/<project>/<run_id>").history() - metrics, training curves, comparisons
EXPERIMENT_LOG.md - full results table with baselines and verdicts
EXPERIMENT_TRACKER.md - check which experiments are done vs still running
Log files - ssh server "tail -100 /path/to/training.log" if no other source
docs/research_contract.md or project notes - intended claims and experiment design

Assemble the key information:

What experiments were run (method, dataset, config)
Main metrics and baseline comparisons (deltas)
The intended claim these experiments were designed to test
Any known confounds or caveats

Step 2: Secondary Codex Judgment

Send the collected results to a secondary Codex agent for objective evaluation:

spawn_agent:
  model: REVIEWER_MODEL
  reasoning_effort: xhigh
  message: |
    RESULT-TO-CLAIM EVALUATION

    I need you to judge whether experimental results support the intended claim.

    Intended claim: [the claim these experiments test]

    Experiments run:
    [list experiments with method, dataset, metrics]

    Results:
    [paste key numbers, comparison deltas, significance]

    Baselines:
    [baseline numbers and sources - reproduced or from paper]

    Known caveats:
    [any confounding factors, limited datasets, missing comparisons]

    Please evaluate:
    1. claim_supported: yes | partial | no
    2. what_results_support: what the data actually shows
    3. what_results_dont_support: where the data falls short of the claim
    4. missing_evidence: specific evidence gaps
    5. suggested_claim_revision: if the claim should be strengthened, weakened, or reframed
    6. next_experiments_needed: specific experiments to fill gaps (if any)
    7. confidence: high | medium | low

    Be honest. Do not inflate claims beyond what the data supports.
    A single positive result on one dataset does not support a general claim.

If delegation is unavailable, run the same evaluation locally and mark the verdict [pending external review] instead of blocking the pipeline.

Step 3: Parse and Normalize

Extract structured fields from the response:

- claim_supported: yes | partial | no
- what_results_support: "..."
- what_results_dont_support: "..."
- missing_evidence: "..."
- suggested_claim_revision: "..."
- next_experiments_needed: "..."
- confidence: high | medium | low

Step 4: Route Based on Verdict

`no` - Claim not supported

Record a postmortem in findings.md:
- What was tested, what failed, and hypotheses for why
- Constraints for future attempts (what not to try again)
Update the project pipeline status in project notes
Decide whether to pivot to the next idea from IDEA_CANDIDATES.md or try an alternative approach

`partial` - Claim partially supported

Update the working claim to reflect what is supported
Record the gap in findings.md
Design and run supplementary experiments to fill evidence gaps
Re-run /result-to-claim after supplementary experiments complete
If the same claim gets multiple partial verdicts, record the analysis in findings.md and consider narrowing the claim scope or switching ideas

`yes` - Claim supported

Record the confirmed claim in project notes
If ablation studies are incomplete, trigger /ablation-planner
If all evidence is in, move to paper writing

Rules

The secondary Codex agent is the judge, not the local executor. The local executor collects evidence and routes; the reviewer agent evaluates. This prevents post-hoc rationalization.
Do not inflate claims beyond what the data supports. If the verdict says partial, do not round up to yes.
A single positive result on one dataset does not support a general claim. Be honest about scope.
If confidence is low, treat the judgment as inconclusive and add experiments rather than committing to a claim.
If reviewer delegation is unavailable, make the best local judgment you can and mark it [pending external review].
Always record the verdict and reasoning in findings.md, regardless of outcome.

Related Skills

shaun-z/paper-illustration-image2

development

VerifiedTrustedCommunity

Generate publication-quality academic illustrations through a local Codex app-server bridge that uses Codex native image generation. This is a separate experimental alternative to `paper-illustration`, intended for Claude Code users who want a GPT-image-style renderer without modifying the original skill.

SKILL.mdUpdated Apr 25, 2026

shaun-z/paper-illustration-image2

shaun-z/overleaf-sync

development

VerifiedTrustedCommunity

Two-way sync between a local paper directory and an Overleaf project via the Overleaf Git bridge (Premium feature). Lets you keep ARIS audit/edit workflows on the local copy while collaborators edit in the Overleaf web UI. Token never touches the agent — user does the one-time auth via macOS Keychain. Use when user says "同步 overleaf", "overleaf sync", "推送到 overleaf", "connect overleaf", "Overleaf 桥接", "pull overleaf", "push overleaf", or wants to bridge their ARIS paper directory with an Overleaf project.

SKILL.mdUpdated Apr 25, 2026

shaun-z/overleaf-sync

shaun-z/citation-audit

development

VerifiedTrustedCommunity

Zero-context verification that every bibliographic entry in the paper is real, correctly attributed, and used in a context the cited paper actually supports. Uses a fresh cross-model reviewer with web/DBLP/arXiv lookup to catch hallucinated authors, wrong years, fabricated venues, version mismatches, and wrong-context citations (cite present but the cited paper does not establish the claim). Use when user says "审查引用", "check citations", "citation audit", "verify references", "引用核对", or before submission to ensure bibliography integrity.

SKILL.mdUpdated Apr 20, 2026

shaun-z/citation-audit

shaun-z/writing-systems-papers

data-ai

VerifiedTrustedCommunity

Paragraph-level structural blueprint for 10-12 page systems papers targeting OSDI, SOSP, ASPLOS, NSDI, and EuroSys. Provides page allocation, paragraph templates, and writing patterns. Use when user says "写系统论文", "systems paper structure", "OSDI paper", "SOSP paper", or wants fine-grained structural guidance for a systems conference submission.

SKILL.mdUpdated Apr 17, 2026

shaun-z/writing-systems-papers

Download

For Claude Desktop. Download once, then upload the file in the app — no terminal needed.

Need help? View full Cowork setup guide →

Install manually

Choose your platform

# Clone the repo
git clone https://github.com/shaun-z/auto-claude-code-research-in-sleep.git

# Copy into Claude Code skills folder (global)
cp -r auto-claude-code-research-in-sleep/skills/skills-codex/result-to-claim ~/.claude/skills/

Claude Code Skills — official skills path docs.

Repository

shaun-z/auto-claude-code-research-in-sleep

Compatible with

Claude Code

OpenAI Codex CLI

ChatGPT

Adoption

shaun-z/result-to-claim

$ install --global

Security Scan Results

SKILL.md

Result-to-Claim Gate

Context: $ARGUMENTS

Constants

When to Use

Workflow

Step 1: Collect Results

Step 2: Secondary Codex Judgment

Step 3: Parse and Normalize

Step 4: Route Based on Verdict

no - Claim not supported

partial - Claim partially supported

yes - Claim supported

Rules

Related Skills

shaun-z/paper-illustration-image2

shaun-z/overleaf-sync

shaun-z/citation-audit

shaun-z/writing-systems-papers

shaun-z/result-to-claim

$ install --global

Security Scan Results

SKILL.md

Result-to-Claim Gate

Context: $ARGUMENTS

Constants

When to Use

Workflow

Step 1: Collect Results

Step 2: Secondary Codex Judgment

Step 3: Parse and Normalize

Step 4: Route Based on Verdict

no - Claim not supported

partial - Claim partially supported

yes - Claim supported

Rules

Related Skills

shaun-z/paper-illustration-image2

shaun-z/overleaf-sync

shaun-z/citation-audit

shaun-z/writing-systems-papers

`no` - Claim not supported

`partial` - Claim partially supported

`yes` - Claim supported

`no` - Claim not supported

`partial` - Claim partially supported

`yes` - Claim supported