Adoption

Agent Skills are supported by leading AI development tools.

VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory VS Code Gemini CLI GitHub Goose Amp Cursor Claude Code Letta OpenCode Claude OpenAI Codex Factory

andyzhuang/claw-ancestry-pca

Name: claw-ancestry-pca
Author: andyzhuang

skills/ancestry-pca/SKILL.md

npx skillsauth add andyzhuang/openlife claw-ancestry-pca

Clean

TrivyContainer and dependency vulnerability scanner

Clean

SemgrepStatic code analysis for vulnerabilities

Clean

mcp-scan (Snyk)Model Context Protocol security validation

Skipped

Snyk (dep)Open source security scanning

Skipped

Socket.devSupply chain security analysis

Skipped

VirusTotalMulti-engine malware detection

Skipped

CrowdStrikeAdvanced threat intelligence

Skipped

OSV-ScannerOpen Source Vulnerability database check

Skipped

OWASP Dep-Check

🦖 Ancestry Decomposition PCA

Place your study cohort in global genetic context by computing a joint PCA against the Simons Genome Diversity Project (SGDP) — 345 samples from 164 populations spanning every inhabited continent.

What it does

Takes your VCF + population map as input
Finds common variants between your cohort and the SGDP reference panel (bundled)
Runs PLINK PCA on the merged dataset
Separates your cohort from SGDP reference samples
Matches SGDP samples to their population labels (164 populations)
Generates a publication-quality multi-panel figure:
- Panel A: PC1 vs PC2 — main population structure of your cohort
- Panel B: PC3 vs PC2 with regional groupings and confidence ellipses
- Panel C: PC3 vs PC1 with language/cultural groupings
- Panel D: Global context — your samples (circles) vs SGDP (triangles)
Produces a markdown report with variance explained, population assignments, and reproducibility bundle

Why this exists

If you ask ChatGPT to "run a PCA against a global reference panel," it will:

Not know which reference panel to use
Hallucinate PLINK flags for merging datasets with different variant sets
Skip IBD removal (related individuals distort PCA)
Not normalise contig names between your VCF and the reference
Produce a single scatter plot with no population labels

This skill encodes the correct methodological decisions:

Uses SGDP (the gold-standard reference for global diversity)
Handles contig normalisation (chr1 vs 1)
Filters to common biallelic SNPs shared between datasets
Removes related individuals via IBD checks
Produces publication-quality multi-panel figures with confidence ellipses
Differentiates your samples (circles) from reference (triangles)

Reference Panel

The skill bundles the SGDP v4 dataset (Mallick et al., 2016, Nature):

345 samples from 164 populations
Whole-genome sequencing at high coverage
MAF > 0.1% filter applied
Populations span: Africa, Americas, Central/South Asia, East Asia, Europe, Middle East, Oceania

Usage

python ancestry_pca.py \
    --vcf your_cohort.vcf.gz \
    --pop-map your_populations.tsv \
    --output ancestry_report

Demo (works out of the box)

python ancestry_pca.py --demo --output demo_report

The demo uses pre-computed PCA results from the Peruvian Genome Project (736 samples, 28 populations) and generates the full 4-panel figure instantly.

Example Output

Ancestry Decomposition PCA
==========================
Cohort: 736 samples, 28 populations
Reference: SGDP (345 samples, 164 populations)
Common variants: 42,831 biallelic SNPs

Variance explained:
  PC1: 51.44%  PC2: 21.70%  PC3: 6.70%

Panel D — Global Context:
  Cohort samples cluster between European and East Asian
  reference populations, with Amazonian groups showing
  distinct positioning from Highland and Coastal groups.

Figures saved to: ancestry_report/
  Figure3_PCA_composite.png (300 dpi)
  Figure3_PCA_composite.pdf (vector)

Reproducibility:
  commands.sh | environment.yml | checksums.sha256

Interpretation Guide

PC1 typically captures the largest axis of global differentiation (often Africa vs non-Africa)
PC2 separates major continental groups (Europe, East Asia, Americas)
PC3 often reveals finer substructure within continental groups
Confidence ellipses show 2.5 standard deviations around each population cluster
Your samples shown as circles, SGDP reference as triangles

Citation

If you use this skill in a publication, please cite:

Mallick, S. et al. (2016). The Simons Genome Diversity Project. Nature, 538, 201-206.
Corpas, M. (2026). OpenLife. https://github.com/OpenLife/OpenLife

andyzhuang/claw-ancestry-pca

skills/ancestry-pca/SKILL.md

Ancestry decomposition PCA against the Simons Genome Diversity Project

26 stars

data-ai

Updated Apr 4, 2026

$ install --global

skillsauth

npx skillsauth add andyzhuang/openlife claw-ancestry-pca

Install this skill globally with one command. Works with Claude Code, Cursor, and Windsurf.

Security Scan Results

3 of 9 scanners reported clean

Some scanners were skipped, did not run, or reported a non-clean status. Review each row below.

Scanners Passed

Scanners in report

Clean

TrivyContainer and dependency vulnerability scanner

95%

Clean

SemgrepStatic code analysis for vulnerabilities

95%

Clean

mcp-scan (Snyk)Model Context Protocol security validation

95%

Skipped

Snyk (dep)Open source security scanning

50%

Skipped

Socket.devSupply chain security analysis

50%

Skipped

VirusTotalMulti-engine malware detection

50%

Skipped

CrowdStrikeAdvanced threat intelligence

50%

Skipped

OSV-ScannerOpen Source Vulnerability database check

50%

Skipped

OWASP Dep-Check

50%

Last scanned: Apr 23, 2026, 1:36 AM3.4s1 file scanned

SKILL.md

name:: claw-ancestry-pca
version:: 0.1.0
description:: Ancestry analysis report with population assignments and statistics
author:: Manuel Corpas
license:: MIT
- name:: report
type:: file
format:: markdown
category:: bioinformatics
homepage:: https://github.com/OpenLife/OpenLife
min_python:: 3.9

🦖 Ancestry Decomposition PCA

Place your study cohort in global genetic context by computing a joint PCA against the Simons Genome Diversity Project (SGDP) — 345 samples from 164 populations spanning every inhabited continent.

What it does

Takes your VCF + population map as input
Finds common variants between your cohort and the SGDP reference panel (bundled)
Runs PLINK PCA on the merged dataset
Separates your cohort from SGDP reference samples
Matches SGDP samples to their population labels (164 populations)
Generates a publication-quality multi-panel figure:
- Panel A: PC1 vs PC2 — main population structure of your cohort
- Panel B: PC3 vs PC2 with regional groupings and confidence ellipses
- Panel C: PC3 vs PC1 with language/cultural groupings
- Panel D: Global context — your samples (circles) vs SGDP (triangles)
Produces a markdown report with variance explained, population assignments, and reproducibility bundle

Why this exists

If you ask ChatGPT to "run a PCA against a global reference panel," it will:

Not know which reference panel to use
Hallucinate PLINK flags for merging datasets with different variant sets
Skip IBD removal (related individuals distort PCA)
Not normalise contig names between your VCF and the reference
Produce a single scatter plot with no population labels

This skill encodes the correct methodological decisions:

Uses SGDP (the gold-standard reference for global diversity)
Handles contig normalisation (chr1 vs 1)
Filters to common biallelic SNPs shared between datasets
Removes related individuals via IBD checks
Produces publication-quality multi-panel figures with confidence ellipses
Differentiates your samples (circles) from reference (triangles)

Reference Panel

The skill bundles the SGDP v4 dataset (Mallick et al., 2016, Nature):

345 samples from 164 populations
Whole-genome sequencing at high coverage
MAF > 0.1% filter applied
Populations span: Africa, Americas, Central/South Asia, East Asia, Europe, Middle East, Oceania

Usage

python ancestry_pca.py \
    --vcf your_cohort.vcf.gz \
    --pop-map your_populations.tsv \
    --output ancestry_report

Demo (works out of the box)

python ancestry_pca.py --demo --output demo_report

The demo uses pre-computed PCA results from the Peruvian Genome Project (736 samples, 28 populations) and generates the full 4-panel figure instantly.

Example Output

Ancestry Decomposition PCA
==========================
Cohort: 736 samples, 28 populations
Reference: SGDP (345 samples, 164 populations)
Common variants: 42,831 biallelic SNPs

Variance explained:
  PC1: 51.44%  PC2: 21.70%  PC3: 6.70%

Panel D — Global Context:
  Cohort samples cluster between European and East Asian
  reference populations, with Amazonian groups showing
  distinct positioning from Highland and Coastal groups.

Figures saved to: ancestry_report/
  Figure3_PCA_composite.png (300 dpi)
  Figure3_PCA_composite.pdf (vector)

Reproducibility:
  commands.sh | environment.yml | checksums.sha256

Interpretation Guide

PC1 typically captures the largest axis of global differentiation (often Africa vs non-Africa)
PC2 separates major continental groups (Europe, East Asia, Americas)
PC3 often reveals finer substructure within continental groups
Confidence ellipses show 2.5 standard deviations around each population cluster
Your samples shown as circles, SGDP reference as triangles

Citation

If you use this skill in a publication, please cite:

Mallick, S. et al. (2016). The Simons Genome Diversity Project. Nature, 538, 201-206.
Corpas, M. (2026). OpenLife. https://github.com/OpenLife/OpenLife

Related Skills

andyzhuang/clinical-trials-search

tools

VerifiedTrustedCommunity

Search ClinicalTrials.gov with natural language queries. Find clinical trials, enrollment, and outcomes using Valyu semantic search.

26SKILL.mdUpdated Apr 4, 2026

andyzhuang/clinical-trials-search

andyzhuang/citation-management

development

VerifiedTrustedCommunity

Comprehensive citation management for academic research. Search Google Scholar and PubMed for papers, extract accurate metadata, validate citations, and generate properly formatted BibTeX entries. This skill should be used when you need to find papers, verify citation information, convert DOIs to BibTeX, or ensure reference accuracy in scientific writing.

26SKILL.mdUpdated Apr 4, 2026

andyzhuang/citation-management

andyzhuang/bioservices

development

VerifiedTrustedCommunity

Unified Python interface to 40+ bioinformatics services. Use when querying multiple databases (UniProt, KEGG, ChEMBL, Reactome) in a single workflow with consistent API. Best for cross-database analysis, ID mapping across services. For quick single-database lookups use gget; for sequence/file manipulation use biopython.

26SKILL.mdUpdated Apr 4, 2026

andyzhuang/bioservices

andyzhuang/biorxiv-search

tools

VerifiedTrustedCommunity

Search bioRxiv biology preprints with natural language queries. Semantic search powered by Valyu.

26SKILL.mdUpdated Apr 4, 2026

andyzhuang/biorxiv-search

Download

For Claude Desktop. Download once, then upload the file in the app — no terminal needed.

Need help? View full Cowork setup guide →

Install manually

Choose your platform

# Clone the repo
git clone https://github.com/andyzhuang/openlife.git

# Copy into Claude Code skills folder (global)
cp -r openlife/skills/ancestry-pca ~/.claude/skills/

Claude Code Skills — official skills path docs.

Repository

andyzhuang/openlife

26 stars

Compatible with

Claude Code

OpenAI Codex CLI

ChatGPT